AI Post Transformers

MiniMax Sparse Attention at Million-Token Scale


Listen Later

This episode explores MiniMax Sparse Attention, a long-context transformer design that aims to preserve dense-model quality at million-token scale while sharply reducing the quadratic compute and memory costs of standard attention. It explains how the method combines Grouped Query Attention with blockwise sparse retrieval: a lightweight Index Branch scores past context in blocks, forces a recent local block to stay visible, selects top-k candidate regions, and then lets a Main Branch run exact softmax attention only inside those chosen blocks. The discussion places the paper alongside Longformer, BigBird, Routing Transformers, MInference, and Native Sparse Attention, arguing that its main contribution is a simpler, more GPU-friendly routing scheme that could make sparse attention practical at deployment time. Listeners would find it interesting because it focuses on the real technical tension behind ultra-long-context models: whether this kind of sparse routing can reliably recover rare distant evidence, or whether it mainly wins through recency bias and careful systems engineering.
Sources:
1. MiniMax Sparse Attention — Xunhao Lai, Weiqi Xu, Yufeng Yang, Qiaorui Chen, Yang Xu, Lunbin Zeng, Xiaolong Li, Haohai Sun, Haichao Zhu, Vito Zhang, Pengyu Zhao, 2026
http://arxiv.org/abs/2606.13392
2. Longformer: The Long-Document Transformer — Iz Beltagy, Matthew E. Peters, Arman Cohan, 2020
https://scholar.google.com/scholar?q=Longformer%3A+The+Long-Document+Transformer
3. Big Bird: Transformers for Longer Sequences — Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Amr Ahmed, 2020
https://scholar.google.com/scholar?q=Big+Bird%3A+Transformers+for+Longer+Sequences
4. Efficient Content-Based Sparse Attention with Routing Transformers — Aurko Roy, Mohammad Saffar, Ashish Vaswani, David Grangier, 2020
https://scholar.google.com/scholar?q=Efficient+Content-Based+Sparse+Attention+with+Routing+Transformers
5. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Wenfeng Liang, Wangding Zeng, 2025
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
6. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie et al., 2025
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
7. Optimizing Mixture of Block Attention — Guangxuan Xiao et al., 2025
https://scholar.google.com/scholar?q=Optimizing+Mixture+of+Block+Attention
8. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang et al., 2024
https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention
9. DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads — Guangxuan Xiao et al., 2024
https://scholar.google.com/scholar?q=DuoAttention%3A+Efficient+Long-Context+LLM+Inference+with+Retrieval+and+Streaming+Heads
10. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision
11. FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference — Dongwei Wang et al., 2025
https://arxiv.org/abs/2508.08256
12. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2024
https://arxiv.org/abs/2412.10319
13. Streaming Video Question-Answering with In-context Video KV-Cache Retrieval — Shangzhe Di et al., 2025
https://arxiv.org/abs/2503.00540
14. Native Hybrid Attention for Efficient Sequence Modeling — Jusen Du et al., 2025
https://arxiv.org/abs/2510.07019
15. Rope to Nope and Back Again: A New Hybrid Attention Strategy — Bowen Yang et al., 2025
https://arxiv.org/abs/2501.18795
16. AI Post Transformers: Optimizing Mixture of Block Attention Through Statistical Theory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-18-optimizing-mixture-of-block-attention-th-214f91.mp3
17. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
18. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
19. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
20. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
21. AI Post Transformers: Ministral 3: Cascade Distillation for Long-Context Multimodal Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-cascade-distillation-for-long-context-mu-0ebd1a.mp3
Interactive Visualization: MiniMax Sparse Attention at Million-Token Scale
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof