AI Post Transformers

DART Speeds Up Speculative LLM Decoding


Listen Later

This episode explores the DART paper as a practical attempt to make speculative decoding deliver real end-to-end speedups for memory-bound LLM inference. It explains how exact draft-and-verify decoding works, why accepted chunk length only matters when the drafter is cheap enough, and how DART differs from Medusa and EAGLE by reusing target-model hidden states to predict several future tokens in parallel with a diffusion-inspired draft stage. The discussion focuses on DART’s mechanics, including multi-layer state reuse, masked future slots, N-gram-guided pruning, and a shifted-logit design that makes the first drafted token especially important because an early mistake invalidates the rest of the chunk. Listeners would find it interesting because it connects model architecture choices to real serving constraints like latency, batching, and GPU efficiency, showing where theoretical decoding gains do and do not survive in production.
Sources:
1. DART: Diffusion-Inspired Speculative Decoding for Fast LLM Inference — Fuliang Liu, Xue Li, Ketai Zhao, Yinxi Gao, Ziyan Zhou, Zhonghui Zhang, Zhibin Wang, Wanchun Dou, Sheng Zhong, Chen Tian, 2026
http://arxiv.org/abs/2601.19278
2. Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding — Hemeng Xia, Zijian Wu, Chunxi Zhang, Yonggan Fu, Haoran Sun, Zhicong Liu, Ping Luo, 2024
https://arxiv.org/abs/2401.07851
3. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://arxiv.org/abs/2211.17192
4. EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2024
https://arxiv.org/abs/2401.15077
5. Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion — Jacob K. Christopher, Brian R. Bartoldson, Tal Ben-Nun, Michael Cardei, Bhavya Kailkhura, Ferdinando Fioretto, 2024
https://arxiv.org/abs/2408.05636
6. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2025
https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test
7. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads — Tianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng, Jason D. Lee, Deming Chen, Tri Dao, 2024
https://scholar.google.com/scholar?q=Medusa%3A+Simple+LLM+Inference+Acceleration+Framework+with+Multiple+Decoding+Heads
8. DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding — Guanghao Li, Zhihui Fu, Min Fang, Qibin Zhao, Ming Tang, Chun Yuan, Jun Wang, 2025
https://scholar.google.com/scholar?q=DiffuSpec%3A+Unlocking+Diffusion+Language+Models+for+Speculative+Decoding
9. SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding — Jameson Sandler, Jacob K. Christopher, Thomas Hartvigsen, Nando Fioretto, 2025
https://scholar.google.com/scholar?q=SpecDiff-2%3A+Scaling+Diffusion+Drafter+Alignment+For+Faster+Speculative+Decoding
10. Speculative Decoding: Performance or Illusion? — Xiaoxuan Liu, Jiaxiang Yu, Jongseok Park, Ion Stoica, Alvin Cheung, 2026
https://scholar.google.com/scholar?q=Speculative+Decoding%3A+Performance+or+Illusion%3F
11. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
12. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
13. AI Post Transformers: InfiniGen for Efficient Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-18-infinigen-for-efficient-long-context-llm-143d77.mp3
14. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof