This episode explores VeriCache, a systems paper that asks whether a large language model can draft tokens using a compressed, lossy KV cache and then verify them against the full cache to recover exactly the same greedy-decoding output. It explains why KV cache size has become a central bottleneck in long-context inference, unpacking token dropping, KV quantization, prefix caching, and the speculative decoding ideas that VeriCache turns into a closed-loop draft-and-verify scheme. The discussion argues that fluency is not enough for real deployments because even a single wrong token can break code, structured JSON, or tool calls, so exact token-by-token agreement is the real standard for safe acceleration. Listeners get a clear picture of where serving stacks such as vLLM, Hugging Face TGI, and NVIDIA already use neighboring optimizations, and why VeriCache’s specific lossless-verification approach is both technically appealing and operationally difficult.
Sources:
1. VeriCache: Turning Lossy KV Cache into Lossless LLM Inference — Jiayi Yao, Samuel Shen, Kuntai Du, Shaoting Feng, Dongjoo Seo, Rui Zhang, Yuyang Huang, Yuhan Liu, Shan Lu, Junchen Jiang, 2026
http://arxiv.org/abs/2605.17613
2. Fast Inference from Transformers via Speculative Decoding — Yaniv Leviathan, Matan Kalman, Yossi Matias, 2023
https://proceedings.mlr.press/v202/leviathan23a.html
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Beidi Chen, et al., 2023
https://proceedings.neurips.cc/paper_files/paper/2023/hash/6ceefa7b15572587b78ecfcebb2827f8-Abstract-Conference.html
4. KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache — Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Beidi Chen, Xia Hu, et al., 2024
https://arxiv.org/abs/2402.02750
5. QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache — Rishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Hooper, Kurt Keutzer, Amir Gholami, et al., 2025
https://arxiv.org/abs/2502.10424
6. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding — Ranajoy Sadhukhan, Jian Chen, Zhuoming Chen, et al., 2024
https://scholar.google.com/scholar?q=MagicDec%3A+Breaking+the+Latency-Throughput+Tradeoff+for+Long+Context+Generation+with+Speculative+Decoding
7. Accelerating Large-Scale Reasoning Model Inference: Self-Speculative Decoding with Sparse Attention — Yilong Zhao, Jiaming Tang, Kan Zhu, Zihao Ye, Chi-Chih Chang, Chaofan Lin, Jongseok Park, Guangxuan Xiao, Mohamed S. Abdelfattah, Mingyu Gao, Baris Kasikci, Song Han, and Ion Stoica, 2025
https://scholar.google.com/scholar?q=Accelerating+Large-Scale+Reasoning+Model+Inference%3A+Self-Speculative+Decoding+with+Sparse+Attention
8. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, and Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
9. ShadowServe: Interference-Free KV Cache Fetching for Distributed Prefix Caching — Xingyu Xiang, Raj Joshi, Yuhan Liu, Jiayi Yao, Chenxingyu Zhao, Junchen Jiang, Yang Zhou, Eddie Kohler, and Minlan Yu, 2025
https://scholar.google.com/scholar?q=ShadowServe%3A+Interference-Free+KV+Cache+Fetching+for+Distributed+Prefix+Caching
10. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yuhan Liu, Yihua Cheng, Jiayi Yao, Yuwei An, Xiaokun Chen, Shaoting Feng, et al., 2025
https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference
11. ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs — Jianlong Lei and Shashikant Ilager, 2026
https://arxiv.org/abs/2603.08727
12. Don't Waste Bits! Adaptive KV-Cache Quantization for Lightweight On-Device LLMs — Sayed Pedram Haeri Boroujeni et al., 2026
https://arxiv.org/abs/2604.04722
13. LongSpec: Long-Context Lossless Speculative Decoding with Efficient Drafting and Verification — Penghui Yang et al., 2025
https://arxiv.org/abs/2502.17421
14. RAPID: Long-Context Inference with Retrieval-Augmented Speculative Decoding — Guanzheng Chen et al., 2025
https://arxiv.org/abs/2502.20330
15. LayerKV: Optimizing Large Language Model Serving with Layer-wise KV Cache Management — Yi Xiong et al., 2024
https://arxiv.org/abs/2410.00428
16. TokenLake: A Unified Segment-level Prefix Cache Pool for Fine-grained Elastic Long-Context LLM Serving — Bingyang Wu et al., 2025
https://arxiv.org/abs/2508.17219
17. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025
https://arxiv.org/abs/2507.07400
18. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025
https://arxiv.org/abs/2503.16525
19. KV Cache Offloading for Context-Intensive Tasks — Andrey Bocharnikov et al., 2026
https://arxiv.org/abs/2604.08426
20. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3
21. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
22. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
23. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
24. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3