AI Post Transformers

FlashFuser and Hopper-Era FFN Kernel Fusion


Listen Later

This episode explores how the FlashFuser paper uses Hopper GPU inter-core communication to push kernel fusion beyond the usual single-SM memory limits, especially for transformer feed-forward networks and gated FFNs. It explains why this matters now: H100-class GPUs have gained compute far faster than memory bandwidth, making activation spills to HBM an increasingly painful bottleneck for workloads that can consume 40 to 60 percent of inference time. The discussion walks through Hopper’s distributed shared memory model and FlashFuser’s core idea of coordinating reduce, shuffle, and multiply patterns across SM clusters so large intermediate activations can stay on chip longer. Listeners would find it interesting because it connects compiler techniques, GPU architecture, and real transformer inference bottlenecks into a concrete argument about when newer hardware may finally make more aggressive fusion worthwhile.
Sources:
1. FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection — Ziyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo, Yijia Diao, Minyi Guo, Jidong Zhai, Yu Feng, Chen Zhang, Anbang Wu, Jingwen Leng, 2025
http://arxiv.org/abs/2512.12949
2. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018
https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning
3. FusionStitching: Boosting Memory Intensive Computations for Deep Learning Workloads — Zhen Zheng, Pengzhan Zhao, Guoping Long, Feiwen Zhu, Kai Zhu, Wenyi Zhao, Lansong Diao, Jun Yang, Wei Lin, 2020
https://scholar.google.com/scholar?q=FusionStitching%3A+Boosting+Memory+Intensive+Computations+for+Deep+Learning+Workloads
4. Operator Fusion in XLA: Analysis and Evaluation — Daniel Snider, Ruofan Liang, 2023
https://scholar.google.com/scholar?q=Operator+Fusion+in+XLA%3A+Analysis+and+Evaluation
5. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness — Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Re, 2022
https://scholar.google.com/scholar?q=FlashAttention%3A+Fast+and+Memory-Efficient+Exact+Attention+with+IO-Awareness
6. Benchmarking and Dissecting the Nvidia Hopper GPU Architecture — Weile Luo, Ruibo Fan, Zeyu Li, Dayou Du, Qiang Wang, Xiaowen Chu, 2024
https://scholar.google.com/scholar?q=Benchmarking+and+Dissecting+the+Nvidia+Hopper+GPU+Architecture
7. A Case Study in CUDA Kernel Fusion: Implementing FlashAttention-2 on NVIDIA Hopper Architecture using the CUTLASS Library — Ganesh Bikshandi, Jay Shah, 2023
https://scholar.google.com/scholar?q=A+Case+Study+in+CUDA+Kernel+Fusion%3A+Implementing+FlashAttention-2+on+NVIDIA+Hopper+Architecture+using+the+CUTLASS+Library
8. Scaling Deep Learning Computation over the Inter-Core Connected Intelligence Processor with T10 — Yiqi Liu, Yuqi Xue, Yu Cheng, Lingxiao Ma, Ziming Miao, Jilong Xue, Jian Huang, 2024
https://scholar.google.com/scholar?q=Scaling+Deep+Learning+Computation+over+the+Inter-Core+Connected+Intelligence+Processor+with+T10
9. FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection — Ziyu Huang, Yangjie Zhou, Zihan Liu, Xinhao Luo, Yijia Diao, Minyi Guo, Jidong Zhai, Yu Feng, Chen Zhang, Anbang Wu, Jingwen Leng, 2025
https://scholar.google.com/scholar?q=FlashFuser%3A+Expanding+the+Scale+of+Kernel+Fusion+for+Compute-Intensive+Operators+via+Inter-Core+Connection
10. Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators Fusion — Size Zheng, Siyuan Chen, Peidi Song, Renze Chen, Xiuhong Li, Shengen Yan, Dahua Lin, Jingwen Leng, Yun Liang, 2023
https://scholar.google.com/scholar?q=Chimera%3A+An+Analytical+Optimizing+Framework+for+Effective+Compute-intensive+Operators+Fusion
11. BOLT: Bridging the Gap between Auto-tuners and Hardware-native Performance — Jiarong Xing, Leyuan Wang, Shang Zhang, Jack Chen, Ang Chen, Yibo Zhu, 2022
https://scholar.google.com/scholar?q=BOLT%3A+Bridging+the+Gap+between+Auto-tuners+and+Hardware-native+Performance
12. MCFuser: High-Performance and Rapid Fusion of Memory-Bound Compute-Intensive Operators — Zheng Zhang, Donglin Yang, Xiaobo Zhou, Dazhao Cheng, 2024
https://scholar.google.com/scholar?q=MCFuser%3A+High-Performance+and+Rapid+Fusion+of+Memory-Bound+Compute-Intensive+Operators
13. Deep Kernel Fusion for Transformers — Zixi Zhang, Zhiwen Mo, Yiren Zhao, Robert Mullins, 2026
https://scholar.google.com/scholar?q=Deep+Kernel+Fusion+for+Transformers
14. Benchmarking thread block cluster — approximate; unclear from snippet, 2023-2026
https://scholar.google.com/scholar?q=Benchmarking+thread+block+cluster
15. ClusterSim: modeling thread block clusters in Hopper GPUs — approximate; unclear from snippet, 2023-2026
https://scholar.google.com/scholar?q=ClusterSim%3A+modeling+thread+block+clusters+in+Hopper+GPUs
16. Analysing and Reducing Costs of Deep Learning Compiler Auto-tuning — approximate; unclear from snippet, 2023-2026
https://scholar.google.com/scholar?q=Analysing+and+Reducing+Costs+of+Deep+Learning+Compiler+Auto-tuning
17. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
18. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
19. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
20. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
21. AI Post Transformers: Caffeine: A Unified FPGA for CNNs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffeine-a-unified-fpga-for-cnns-e8acbe.mp3
22. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
Interactive Visualization: FlashFuser and Hopper-Era FFN Kernel Fusion
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof