This episode explores SageAttention2, an ICML 2025 paper on making exact transformer attention faster without changing the underlying computation, focusing on why long-context models still pay a steep quadratic cost and why exact kernels remain important despite sparse and linear alternatives. It explains the paper’s central claim that aggressive low-precision attention can work only with careful numerical repair: queries and keys are pushed to INT4, attention-weight and value computation moves toward FP8, and outlier-smoothing ideas inspired by SmoothQuant are used to keep softmax-sensitive logits from collapsing. The discussion highlights the paper’s most concrete systems contribution, per-thread INT4 quantization aligned to GPU thread fragments and PTX `mma` execution, which aims to get fine-grained scaling without losing the performance win to dequantization overhead. A listener would find it interesting because the episode turns a seemingly narrow kernel optimization into a broader argument about hardware-software co-design, showing how much engineering is required to make lower-bit attention practical rather than just theoretically faster.
Sources:
1. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization — Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen, 2024
http://arxiv.org/abs/2411.10958
2. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning — Tri Dao, 2023
https://scholar.google.com/scholar?q=FlashAttention-2%3A+Faster+Attention+with+Better+Parallelism+and+Work+Partitioning
3. INT-FlashAttention: Enabling Flash Attention for INT8 Quantization — Shimao Chen, Zirui Liu, Zhiying Wu, et al., 2024
https://scholar.google.com/scholar?q=INT-FlashAttention%3A+Enabling+Flash+Attention+for+INT8+Quantization
4. SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference Acceleration — Jintao Zhang, Jia Wei, Haofeng Huang, Pengle Zhang, Jun Zhu, Jianfei Chen, 2025
https://scholar.google.com/scholar?q=SageAttention%3A+Accurate+8-Bit+Attention+for+Plug-and-play+Inference+Acceleration
5. SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 Quantization — Jintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei, Jun Zhu, Jianfei Chen, 2025
https://scholar.google.com/scholar?q=SageAttention2%3A+Efficient+Attention+with+Thorough+Outlier+Smoothing+and+Per-thread+INT4+Quantization
6. Understanding and Overcoming the Challenges of Efficient Transformer Quantization — Yelysei Bondarenko, Markus Nagel, Tijmen Blankevoort, 2021
https://scholar.google.com/scholar?q=Understanding+and+Overcoming+the+Challenges+of+Efficient+Transformer+Quantization
7. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han, 2023
https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models
8. Outlier Suppression+: Accurate quantization of large language models by equivalent and optimal shifting and scaling — Xiuying Wei, Yunchen Zhang, Yuhang Li, Xiangguo Zhang, Ruihao Gong, Jinyang Guo, Xianglong Liu, 2023
https://scholar.google.com/scholar?q=Outlier+Suppression%2B%3A+Accurate+quantization+of+large+language+models+by+equivalent+and+optimal+shifting+and+scaling
9. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs — Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, et al., 2024
https://scholar.google.com/scholar?q=QuaRot%3A+Outlier-Free+4-Bit+Inference+in+Rotated+LLMs
10. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision — Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, Tri Dao, 2024
https://scholar.google.com/scholar?q=FlashAttention-3%3A+Fast+and+Accurate+Attention+with+Asynchrony+and+Low-precision
11. QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving — Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han, 2024
https://scholar.google.com/scholar?q=QServe%3A+W4A8KV4+Quantization+and+System+Co-design+for+Efficient+LLM+Serving
12. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang et al., 2024
https://arxiv.org/abs/2407.02490
13. SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention — Qianchao Zhu et al., 2024
https://arxiv.org/abs/2406.15486
14. FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference — Xunhao Lai et al., 2025
https://arxiv.org/abs/2502.20766
15. Activation Outliers in Transformer Quantization: Reproduction, Statistical Analysis, and Deployment Tradeoffs — Pranav Kumar Kaliaperumal, 2026
https://arxiv.org/abs/2603.04308
16. BAPS: A Fine-Grained Low-Precision Scheme for Softmax in Attention via Block-Aware Precision reScaling — Zisheng Ye et al., 2026
https://arxiv.org/abs/2602.02071
17. Softpick: No Attention Sink, No Massive Activations with Rectified Softmax — Zayd M. K. Zuhri et al., 2025
https://arxiv.org/abs/2504.20966
18. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
19. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
20. AI Post Transformers: NanoFlow and the Future of LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-nanoflow-and-the-future-of-llm-serving-7429c9.mp3
Interactive Visualization: SageAttention2 and Fast Exact INT4 Attention