This episode explores the ACE proposal from AMD, Intel, and the x86 Ecosystem Advisory Group, which would add matrix-native AI instructions to x86 CPUs so transformer workloads can run dense linear algebra more efficiently without changing the models themselves. It explains why AVX10 and VNNI fall short for GEMM-heavy inference, introducing outer-product updates, 2-D tile registers, and the reuse of the AMX palette model so operating systems and compilers can handle the new state within familiar x86 mechanisms. The discussion also challenges the proposal’s headline 16x compute-density claim for INT8 and BF16, arguing that real speed depends on full-kernel costs like packing, memory traffic, conversions, tails, and cache behavior. It also examines OCP FP8, MXFP8, MXINT8, and BF16 support as a sign that low-precision AI now depends on tight coordination between ISA design, quantization rules, and kernel implementation, making the proposal interesting both technically and strategically.
Sources:
1. ACE: Matrix-Native AI Extensions for x86
https://x86ecosystem.org/wp-content/uploads/2026/03/ACE-Whitepaper-v1.pdf
2. The AI Compute Extensions (ACE) for x86 — Stuart Biles, Brian Thompto, Michael Estlick, Eric Schwarz, Thomas Fox, Gabriel Loh, Marius Evers, Michael Clark, Alexander Heinecke, Pradeep Dubey, Ido Ouziel, 2026
https://scholar.google.com/scholar?q=The+AI+Compute+Extensions+%28ACE%29+for+x86
3. A matrix math facility for Power ISA(TM) processors — José E. Moreira, Kit Barton, Steven Battle, Peter Bergner, Ramon Bertran and others, 2021
https://scholar.google.com/scholar?q=A+matrix+math+facility+for+Power+ISA%28TM%29+processors
4. Hello SME! Generating Fast Matrix Multiplication Kernels Using the Scalable Matrix Extension — Stefan Remke, Alexander Breuer, 2024
https://scholar.google.com/scholar?q=Hello+SME%21+Generating+Fast+Matrix+Multiplication+Kernels+Using+the+Scalable+Matrix+Extension
5. SparAMX: Accelerating Compressed LLMs Token Generation on AMX-powered CPUs — Ahmed F. AbouElhamayed, Jordan Dotzel, Yash Akhauri, Chi-Chih Chang, Sameh Gobriel, J. Pablo Muñoz, Vui Seng Chua, Nilesh Jain, Mohamed S. Abdelfattah, 2025
https://scholar.google.com/scholar?q=SparAMX%3A+Accelerating+Compressed+LLMs+Token+Generation+on+AMX-powered+CPUs
6. Automating the Last-Mile for High Performance Dense Linear Algebra — Richard Michael Veras, Tze Meng Low, Tyler Michael Smith, Robert A. van de Geijn, Franz Franchetti, 2017
https://scholar.google.com/scholar?q=Automating+the+Last-Mile+for+High+Performance+Dense+Linear+Algebra
7. Microscaling Data Formats for Deep Learning — Bita Darvish Rouhani et al., 2023
https://scholar.google.com/scholar?q=Microscaling+Data+Formats+for+Deep+Learning
8. Recipes for Pre-training LLMs with MXFP8 — Asit Mishra, Dusan Stosic, Simon Layton, 2025
https://scholar.google.com/scholar?q=Recipes+for+Pre-training+LLMs+with+MXFP8
9. Fast Matrix Multiplication via Compiler-only Layered Data Reorganization and Intrinsic Lowering — Braedy Kuzma et al., 2023
https://scholar.google.com/scholar?q=Fast+Matrix+Multiplication+via+Compiler-only+Layered+Data+Reorganization+and+Intrinsic+Lowering
10. THOR: A Non-Speculative Value Dependent Timing Side Channel Attack Exploiting Intel AMX — Farshad Dizani et al., 2025
https://scholar.google.com/scholar?q=THOR%3A+A+Non-Speculative+Value+Dependent+Timing+Side+Channel+Attack+Exploiting+Intel+AMX
11. Compute or Load KV Cache? Why Not Both? (https://arxiv.org/abs/2410.03065) — Shuowei Jin et al., 2024
https://scholar.google.com/scholar?q=Compute+or+Load+KV+Cache%3F+Why+Not+Both%3F+%28https%3A%2F%2Farxiv.org%2Fabs%2F2410.03065%29
12. SOLE: Hardware-Software Co-design of Softmax and LayerNorm for Efficient Transformer Inference (https://arxiv.org/abs/2510.17189) — Wenxun Wang et al., 2025
https://scholar.google.com/scholar?q=SOLE%3A+Hardware-Software+Co-design+of+Softmax+and+LayerNorm+for+Efficient+Transformer+Inference+%28https%3A%2F%2Farxiv.org%2Fabs%2F2510.17189%29
13. Towards Fully FP8 GEMM LLM Training at Scale (https://arxiv.org/abs/2505.20524) — Alejandro Hernandez-Cano et al., 2025
https://scholar.google.com/scholar?q=Towards+Fully+FP8+GEMM+LLM+Training+at+Scale+%28https%3A%2F%2Farxiv.org%2Fabs%2F2505.20524%29
14. Pretraining Large Language Models with NVFP4 (https://arxiv.org/abs/2509.25149) — Felix Abecassis et al. (NVIDIA), 2025
https://scholar.google.com/scholar?q=Pretraining+Large+Language+Models+with+NVFP4+%28https%3A%2F%2Farxiv.org%2Fabs%2F2509.25149%29
15. SageAttention3: Microscaling FP4 Attention for Inference and An Exploration of 8-Bit Training (https://arxiv.org/abs/2505.11594) — Jintao Zhang et al., 2025
https://scholar.google.com/scholar?q=SageAttention3%3A+Microscaling+FP4+Attention+for+Inference+and+An+Exploration+of+8-Bit+Training+%28https%3A%2F%2Farxiv.org%2Fabs%2F2505.11594%29
16. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
17. AI Post Transformers: LPU Chip for Low-Latency LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-lpu-chip-for-low-latency-llm-inference-be13c3.mp3
18. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
19. AI Post Transformers: FlatAttention for Tile-Based Accelerator Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-flatattention-for-tile-based-accelerator-56e6ca.mp3
Interactive Visualization: ACE: Matrix-Native AI Extensions for x86