AI Post Transformers

DSpark Improves Speculative Decoding Acceptance Rates


Listen Later

This episode explores DSpark, a DeepSeek-AI paper on improving speculative decoding by starting from a DFlash-style block-parallel draft model and increasing how often a larger verifier accepts its proposed tokens. It explains the mechanics of speculative decoding in plain language, situates DSpark within earlier blockwise and multi-token prediction work, and notes that the technique is already used in serving stacks such as vLLM, TensorRT-LLM, and SGLang. The discussion focuses on DSpark’s concrete additions: a Markov head that feeds previous-token information into draft logits, a confidence head that estimates whether drafted tokens will survive verification, and a training recipe centered on knowledge distillation. It is interesting because it treats inference speed as an operational systems problem, arguing that higher acceptance matters but only alongside draft latency, verifier cost, batching, and scheduler behavior.
Sources:
1. DSpark Improves Speculative Decoding Acceptance Rates
https://github.com/deepseek-ai/DeepSpec/blob/main/DSpark_paper.pdf
2. DFlash: Block Diffusion for Flash Speculative Decoding — Jian Chen, Yesheng Liang, Zhijian Liu, 2026
https://scholar.google.com/scholar?q=DFlash%3A+Block+Diffusion+for+Flash+Speculative+Decoding
3. EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test — Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang, 2025
https://scholar.google.com/scholar?q=EAGLE-3%3A+Scaling+up+Inference+Acceleration+of+Large+Language+Models+via+Training-Time+Test
4. Decoding Speculative Decoding — Minghao Yan, Saurabh Agarwal, Shivaram Venkataraman, 2024
https://scholar.google.com/scholar?q=Decoding+Speculative+Decoding
5. Speculative Decoding with a Speculative Vocabulary — Miles Williams, Young D. Kwon, Rui Li, Alexandros Kouris, Stylianos I. Venieris, 2026
https://scholar.google.com/scholar?q=Speculative+Decoding+with+a+Speculative+Vocabulary
6. DFlare: Scaling Up Draft Capacity for Block Diffusion Speculative Decoding — Jiebin Zhang, Zhenghan Yu, Song Liu, Eugene J. Yu, Zheng Li, Dawei Zhu, Jiangshan Duo, Weimin Xiong, Yifan Song, Guanghua Yu, Jianchen Zhu, Sujian Li, 2026
https://scholar.google.com/scholar?q=DFlare%3A+Scaling+Up+Draft+Capacity+for+Block+Diffusion+Speculative+Decoding
7. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
8. AI Post Transformers: Adaptive Control for Batched Speculative Decoding in LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/adaptive-control-for-batched-speculative-decoding-in-llm-serving/
9. AI Post Transformers: JETSPEC and Parallel Tree Speculative Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-27-jetspec-and-parallel-tree-speculative-de-3d144c.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof