AI Post Transformers

Efficient Post-Training Quantization with FP8


Listen Later

This episode explores how post-training quantization can convert already-trained models into 8-bit floating point formats for cheaper inference, and why FP8 may outperform the older INT8 approach on modern transformers, LLMs, and diffusion models. It explains the tradeoff between exponent range and mantissa precision across FP8 formats such as E4M3, E5M2, and E3M4, with particular attention to how FP8 handles activation outliers and dynamic range more gracefully than fixed-scale INT8. The discussion centers on a hardware-aware deployment recipe, including which operators can stay quantized, where higher-precision accumulation still matters, and how BatchNorm recalibration helps low-precision inference match full-precision behavior. Listeners get a concrete result: across 75 architectures and more than 200 task cases, the paper reports 92.64% workload coverage for FP8 versus 65.87% for INT8, with E4M3 looking strongest for NLP while E3M4 is slightly better for some vision workloads.
Sources:
1. Efficient Post-training Quantization with FP8 Formats — Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, Mengni Wang, 2023
http://arxiv.org/abs/2309.14592
2. Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference — Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, et al., 2018
https://scholar.google.com/scholar?q=Quantization+and+Training+of+Neural+Networks+for+Efficient+Integer-Arithmetic-Only+Inference
3. A White Paper on Neural Network Quantization — Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, Tijmen Blankevoort, 2021
https://scholar.google.com/scholar?q=A+White+Paper+on+Neural+Network+Quantization
4. LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale — Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer, 2022
https://scholar.google.com/scholar?q=LLM.int8%28%29%3A+8-bit+Matrix+Multiplication+for+Transformers+at+Scale
5. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models — Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han, 2022
https://scholar.google.com/scholar?q=SmoothQuant%3A+Accurate+and+Efficient+Post-Training+Quantization+for+Large+Language+Models
6. 8-bit Numerical Formats for Deep Neural Networks — Badreddine Noune, Philip Jones, Daniel Justus, Dominic Masters, Carlo Luschi, 2022
https://scholar.google.com/scholar?q=8-bit+Numerical+Formats+for+Deep+Neural+Networks
7. FP8 Formats for Deep Learning — Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, et al., 2022
https://scholar.google.com/scholar?q=FP8+Formats+for+Deep+Learning
8. FP8 Quantization: The Power of the Exponent — Andrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel, Jorn Peters, Tijmen Blankevoort, 2022
https://scholar.google.com/scholar?q=FP8+Quantization%3A+The+Power+of+the+Exponent
9. Efficient Post-training Quantization with FP8 Formats — Haihao Shen, Naveen Mellempudi, Xin He, Qun Gao, Chang Wang, Mengni Wang, 2024
https://scholar.google.com/scholar?q=Efficient+Post-training+Quantization+with+FP8+Formats
10. Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks — Xiaohu Sun et al., 2019
https://scholar.google.com/scholar?q=Hybrid+8-bit+Floating+Point+%28HFP8%29+Training+and+Inference+for+Deep+Neural+Networks
11. Outlier Suppression: Pushing the Limit of Low-bit Transformer Language Models — Xiuying Wei et al., 2022
https://scholar.google.com/scholar?q=Outlier+Suppression%3A+Pushing+the+Limit+of+Low-bit+Transformer+Language+Models
12. Quantizable transformers: Removing outliers by helping attention heads do nothing — Bondarenko et al. (approx.), 2023?
https://scholar.google.com/scholar?q=Quantizable+transformers%3A+Removing+outliers+by+helping+attention+heads+do+nothing
13. Understanding and minimising outlier features in transformer training — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=Understanding+and+minimising+outlier+features+in+transformer+training
14. QuanTool: A Benchmarking Framework for Evaluating Post-Training Quantization with Best Practices for Transformer Models — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=QuanTool%3A+A+Benchmarking+Framework+for+Evaluating+Post-Training+Quantization+with+Best+Practices+for+Transformer+Models
15. GO-ViT: Fully Quantizing Vision Transformers by Grouping Outlier Channels — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=GO-ViT%3A+Fully+Quantizing+Vision+Transformers+by+Grouping+Outlier+Channels
16. Understanding int4 quantization for language models: latency speedup, composability, and failure cases — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=Understanding+int4+quantization+for+language+models%3A+latency+speedup%2C+composability%2C+and+failure+cases
17. Kvquant: Towards 10 million context length llm inference with kv cache quantization — author list not verified from provided snippet, 2024?
https://scholar.google.com/scholar?q=Kvquant%3A+Towards+10+million+context+length+llm+inference+with+kv+cache+quantization
18. AI Post Transformers: MIOpen and AMD's Open Deep Learning Primitives — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-23-miopen-and-amds-open-deep-learning-primi-052f82.mp3
19. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
20. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
Interactive Visualization: Efficient Post-Training Quantization with FP8
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof