AI Post Transformers

MIOpen and AMD's Open Deep Learning Primitives


Listen Later

This episode explores AMD’s open-source MIOpen library and why deep learning primitives such as convolution, pooling, normalization, and activations are the layer where model performance meets GPU hardware reality. It explains how CNN throughput depends on low-level execution choices, comparing approaches such as im2col-plus-GEMM and Winograd convolution, and shows why libraries like MIOpen need solver-based algorithm selection and auto-tuning to match different workload shapes, precisions, and GPUs. The discussion also covers mixed-precision support, especially bfloat16, along with kernel fusion and composable kernels as ways to reduce memory traffic and launch overhead while keeping vendor-library speed. Listeners would find it interesting because it turns “invisible infrastructure” into a concrete systems story about how open-source GPU software can shape real model training and inference performance.
Sources:
1. MIOpen: An Open Source Library For Deep Learning Primitives — Jehandad Khan, Paul Fultz, Artem Tamazov, Daniel Lowell, Chao Liu, Michael Melesse, Murali Nandhimandalam, Kamil Nasyrov, Ilya Perminov, Tejash Shah, Vasilii Filippov, Jing Zhang, Jing Zhou, Bragadeesh Natarajan, Mayank Daga, 2019
http://arxiv.org/abs/1910.00078
2. cuDNN: Efficient Primitives for Deep Learning — Sharan Chetlur, Cliff Woolley, Philippe Vandermersch, Jonathan Cohen, John Tran, Bryan Catanzaro, Evan Shelhamer, 2014
https://scholar.google.com/scholar?q=cuDNN%3A+Efficient+Primitives+for+Deep+Learning
3. MIOpen: An Open Source Library For Deep Learning Primitives — Jehandad Khan, Paul Fultz, Artem Tamazov, Daniel Lowell, Chao Liu, Michael Melesse, Murali Nandhimandalam, Kamil Nasyrov, Ilya Perminov, Tejash Shah, Vasilii Filippov, Jing Zhang, Jing Zhou, Bragadeesh Natarajan, Mayank Daga, 2019
https://scholar.google.com/scholar?q=MIOpen%3A+An+Open+Source+Library+For+Deep+Learning+Primitives
4. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning — Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Meghan Cowan, Haichen Shen, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, Arvind Krishnamurthy, 2018
https://scholar.google.com/scholar?q=TVM%3A+An+Automated+End-to-End+Optimizing+Compiler+for+Deep+Learning
5. Ansor: Generating High-Performance Tensor Programs for Deep Learning — Lianmin Zheng, Chengfan Jia, Minmin Sun, Zhao Wu, Cody Hao Yu, Ameer Haj-Ali, Yida Wang, Jun Yang, Danyang Zhuo, Koushik Sen, Joseph E. Gonzalez, Ion Stoica, 2020
https://scholar.google.com/scholar?q=Ansor%3A+Generating+High-Performance+Tensor+Programs+for+Deep+Learning
6. Automatically tuned linear algebra software — R. Clint Whaley and Jack J. Dongarra, 1998
https://scholar.google.com/scholar?q=Automatically+tuned+linear+algebra+software
7. Fast Algorithms for Convolutional Neural Networks — Andrew Lavin and Scott Gray, 2015
https://scholar.google.com/scholar?q=Fast+Algorithms+for+Convolutional+Neural+Networks
8. TVM: end-to-end optimization stack for deep learning — Tianqi Chen et al., 2018
https://scholar.google.com/scholar?q=TVM%3A+end-to-end+optimization+stack+for+deep+learning
9. Tensor Comprehensions: Framework-Agnostic High-Performance Machine Learning Abstractions — Nicolas Vasilache et al., 2018
https://scholar.google.com/scholar?q=Tensor+Comprehensions%3A+Framework-Agnostic+High-Performance+Machine+Learning+Abstractions
10. oneDNN Graph Compiler: A Hybrid Approach for High-Performance Deep Learning Compilation — approx. Intel oneDNN graph compiler authors, 2023
https://scholar.google.com/scholar?q=oneDNN+Graph+Compiler%3A+A+Hybrid+Approach+for+High-Performance+Deep+Learning+Compilation
11. SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning — approx. TVM / SparseTIR authors, 2023
https://scholar.google.com/scholar?q=SparseTIR%3A+Composable+Abstractions+for+Sparse+Compilation+in+Deep+Learning
12. Autotuning Convolutions Is Easier Than You Think — approx. tensor-compiler autotuning authors, 2023
https://scholar.google.com/scholar?q=Autotuning+Convolutions+Is+Easier+Than+You+Think
13. Haotuner: A Hardware Adaptive Operator Auto-Tuner for Dynamic Shape Tensor Compilers — approx. Haotuner authors, 2024
https://scholar.google.com/scholar?q=Haotuner%3A+A+Hardware+Adaptive+Operator+Auto-Tuner+for+Dynamic+Shape+Tensor+Compilers
14. The Case for Training Large Models in Low Precision — approx. low-precision training authors, 2024
https://scholar.google.com/scholar?q=The+Case+for+Training+Large+Models+in+Low+Precision
15. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
16. AI Post Transformers: FlashFuser and Hopper-Era FFN Kernel Fusion — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-flashfuser-and-hopper-era-ffn-kernel-fus-e1fce9.mp3
17. AI Post Transformers: Automating DNN Compilation for FPGA Accelerators — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-automating-dnn-compilation-for-fpga-acce-6ef9bf.mp3
18. AI Post Transformers: Caffeine: A Unified FPGA for CNNs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-caffeine-a-unified-fpga-for-cnns-e8acbe.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof