This episode explores the paper Is Finer Better? The Limits of Microscaling Formats in Large Language Models and examines why shrinking microscaling block sizes can unexpectedly make low-bit LLM quantization worse instead of better. It walks through how microscaling pairs FP4 weights or activations with shared local FP8 scales, contrasts that setup with coarser quantization schemes, and places the work in the broader move from BF16 and FP8 toward cheaper, more hardware-friendly inference. The central argument is that smaller blocks do reduce element quantization error, but once the shared scale is itself quantized into a limited format like FP8 UE4M3, scale error can dominate and degrade perplexity. Listeners would find it interesting because the discussion turns a seemingly obvious engineering intuition on its head and shows that the real bottleneck in low-bit inference may be the precision of the scaling rule, not just the precision of the values being scaled.
Sources:
1. Is Finer Better? The Limits of Microscaling Formats in Large Language Models — Andrea Fasoli, Monodeep Kar, Chi-Chun Liu, Swagath Venkataramani, Viji Srinivasan, Leland Chang, Naigang Wang, 2026
http://arxiv.org/abs/2601.19026
2. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, and others, 2022
https://arxiv.org/abs/2209.05433
3. With Shared Microexponents, A Little Shifting Goes a Long Way — Bita Rouhani, Ritchie Zhao, Venmugil Elango, Rasoul Shafipour, Mathew Hall, Maral Mesmakhosroshahi, Ankit More, and others, 2023
https://arxiv.org/abs/2302.08007
4. Microscaling Data Formats for Deep Learning — Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, and others, 2023
https://arxiv.org/abs/2310.10537
5. Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization — Vage Egiazarian, Roberto L. Castro, Denis Kuznedelev, Andrei Panferov, Eldar Kurtic, Shubhra Pandit, Alexandre Marques, Mark Kurtz, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh, 2025
https://arxiv.org/abs/2509.23202
6. AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference — Janghwan Lee et al., 2024
https://scholar.google.com/scholar?q=AMXFP4%3A+Taming+Activation+Outliers+with+Asymmetric+Microscaling+Floating-Point+for+4-bit+LLM+Inference
7. Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models — Yun-Chen Lo, Gu-Yeon Wei, David Brooks, 2024
https://scholar.google.com/scholar?q=Nanoscaling+Floating-Point+%28NxFP%29%3A+NanoMantissa%2C+Adaptive+Microexponents%2C+and+Code+Recycling+for+Direct-Cast+Compression+of+Large+Language+Models
8. Elucidating the Design Space of FP4 Training — Robert Hu, Carlo Luschi, Paul Balanca, 2025
https://scholar.google.com/scholar?q=Elucidating+the+Design+Space+of+FP4+Training
9. Finer is Better (with the Right Scaling) — Clemens Schaefer, Gil Tabak, 2026
https://scholar.google.com/scholar?q=Finer+is+Better+%28with+the+Right+Scaling%29
10. Adaptive Block-Scaled Data Types — Jack Cook et al., 2026
https://arxiv.org/abs/2603.28765
11. Diagnosing FP4 inference: a layer-wise and block-wise sensitivity analysis of NVFP4 and MXFP4 — Musa Cim et al., 2026
https://arxiv.org/abs/2603.08747
12. Pretraining large language models with MXFP4 — Musa Cim et al., 2026
https://arxiv.org/abs/2605.09825
13. Dissecting Outlier Dynamics in LLM NVFP4 Pretraining — Peijie Dong et al., 2026
https://arxiv.org/abs/2602.02047
14. DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization — Haokun Lin et al., 2026
https://arxiv.org/abs/2604.17789
15. AdaHOP: Fast and Accurate Low-Precision Training via Outlier-Pattern-Aware Rotation — Seonggon Kim et al., 2026
https://arxiv.org/abs/2604.02525
16. AI Post Transformers: Nemotron 3 Ultra for Long-Horizon Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-17-nemotron-3-ultra-for-long-horizon-agents-32e4a5.mp3
17. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
18. AI Post Transformers: FlashAttention-4 Conquers Asymmetric GPU Hardware Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-06-flashattention-4-conquers-asymmetric-gpu-78839b.mp3