AI Post Transformers

DeltaKV: Compressing KV Caches for Long Context


Listen Later

This episode explores DeltaKV, a method for reducing the huge GPU memory burden of KV caches in long-context language model inference without simply discarding old tokens. It contrasts three strategies for handling long contexts: token eviction, dynamic sparse attention, and true compression, arguing that the cache contains structured redundancy that can be exploited rather than treated as disposable overhead. The discussion highlights DeltaKV’s core idea of keeping a small uncompressed reference set and storing other cache entries as compressed residuals relative to similar past states, drawing an analogy to delta encoding or version control. Listeners would find it interesting because it connects transformer internals, systems constraints, and practical serving performance, including claims of cutting memory to 29 percent of baseline and reaching up to 2x throughput with supporting infrastructure like Sparse-vLLM.
Sources:
1. DeltaKV: Residual-Based KV Cache Compression via Long-Range Similarity — Jitai Hao, Qiang Huang, Yaowei Wang, Min Zhang, Jun Yu, 2026
http://arxiv.org/abs/2602.08005
2. Generalization through Memorization: Nearest Neighbor Language Models — Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, Mike Lewis, 2019
https://scholar.google.com/scholar?q=Generalization+through+Memorization%3A+Nearest+Neighbor+Language+Models
3. Reformer: The Efficient Transformer — Nikita Kitaev, Lukasz Kaiser, Anselm Levskaya, 2020
https://scholar.google.com/scholar?q=Reformer%3A+The+Efficient+Transformer
4. Improving Language Models by Retrieving from Trillions of Tokens — Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann and many others, 2021
https://scholar.google.com/scholar?q=Improving+Language+Models+by+Retrieving+from+Trillions+of+Tokens
5. Memorizing Transformers — Yuhuai Wu, Markus N. Rabe, DeLesley Hutchins, Christian Szegedy, 2022
https://scholar.google.com/scholar?q=Memorizing+Transformers
6. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu, Hanchen Li, Yihua Cheng and others, 2023
https://scholar.google.com/scholar?q=CacheGen%3A+KV+Cache+Compression+and+Streaming+for+Fast+Large+Language+Model+Serving
7. Palu: Compressing KV-Cache with Low-Rank Projection — Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin and others, 2024
https://scholar.google.com/scholar?q=Palu%3A+Compressing+KV-Cache+with+Low-Rank+Projection
8. Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries — Junhyuck Kim, Jongho Park, Jaewoong Cho, Dimitris Papailiopoulos, 2025
https://scholar.google.com/scholar?q=Lexico%3A+Extreme+KV+Cache+Compression+via+Sparse+Coding+over+Universal+Dictionaries
9. Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models — Alina Shutova, Vladimir Malinovskii, Vage Egiazarian and others, 2025
https://scholar.google.com/scholar?q=Cache+Me+If+You+Must%3A+Adaptive+Key-Value+Quantization+for+Large+Language+Models
10. OmniKV: Dynamic Context Selection for Efficient Long-Context LLMs — Jitai Hao, Yuke Zhu, Tian Wang, Jun Yu, Xin Xin, Bo Zheng, Zhaochun Ren, Sheng Guo, 2025
https://scholar.google.com/scholar?q=OmniKV%3A+Dynamic+Context+Selection+for+Efficient+Long-Context+LLMs
11. QUEST: Query-Aware Sparsity for Efficient Long-Context LLM Inference — Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, Song Han, 2024
https://scholar.google.com/scholar?q=QUEST%3A+Query-Aware+Sparsity+for+Efficient+Long-Context+LLM+Inference
12. Palu: KV-Cache Compression with Low-Rank Projection — Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Chong-Yan Chen, Yu-Fang Hu, Pei-Shuo Wang, Ning-Chi Huang, Luis Ceze, Mohamed S. Abdelfattah, Kai-Chiang Wu, 2025
https://scholar.google.com/scholar?q=Palu%3A+KV-Cache+Compression+with+Low-Rank+Projection
13. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li, Huiqiang Jiang, Qianhui Wu, Xufang Luo, Surin Ahn, Chengruidong Zhang, Amir H. Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, Lili Qiu, 2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
14. The Pitfalls of KV Cache Compression — Alex Chen, Renato Geh, Aditya Grover, Guy Van den Broeck, Daniel Israel, 2025
https://scholar.google.com/scholar?q=The+Pitfalls+of+KV+Cache+Compression
15. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
16. Kvlink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. recent LLM systems authors, 2025/2026
https://scholar.google.com/scholar?q=Kvlink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
17. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — approx. recent RAG systems authors, 2025/2026
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
18. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — approx. recent RAG/serving authors, 2025/2026
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
19. KVSink: Understanding and Enhancing the Preservation of Attention Sinks in KV Cache Quantization for LLMs — approx. recent KV quantization authors, 2025/2026
https://scholar.google.com/scholar?q=KVSink%3A+Understanding+and+Enhancing+the+Preservation+of+Attention+Sinks+in+KV+Cache+Quantization+for+LLMs
20. LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference — approx. recent long-context inference authors, 2025/2026
https://scholar.google.com/scholar?q=LLMs+Know+What+to+Drop%3A+Self-Attention+Guided+KV+Cache+Eviction+for+Efficient+Long-Context+Inference
21. Task-KV: Task-Aware KV Cache Optimization via Semantic Differentiation of Attention Heads — approx. recent attention/KV optimization authors, 2025/2026
https://scholar.google.com/scholar?q=Task-KV%3A+Task-Aware+KV+Cache+Optimization+via+Semantic+Differentiation+of+Attention+Heads
22. WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More — approx. recent quantization authors, 2025/2026
https://scholar.google.com/scholar?q=WKVQuant%3A+Quantizing+Weight+and+Key%2FValue+Cache+for+Large+Language+Models+Gains+More
23. AlignedKV: Reducing Memory Access of KV-Cache with Precision-Aligned Quantization — approx. recent KV systems/quantization authors, 2025/2026
https://scholar.google.com/scholar?q=AlignedKV%3A+Reducing+Memory+Access+of+KV-Cache+with+Precision-Aligned+Quantization
24. Making Every Head Count: Sparse Attention Without the Speed-Performance Trade-off — approx. recent sparse attention authors, 2025/2026
https://scholar.google.com/scholar?q=Making+Every+Head+Count%3A+Sparse+Attention+Without+the+Speed-Performance+Trade-off
25. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention — approx. recent sparse systems authors, 2025/2026
https://scholar.google.com/scholar?q=Native+Sparse+Attention%3A+Hardware-Aligned+and+Natively+Trainable+Sparse+Attention
26. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3
27. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
28. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
29. AI Post Transformers: Native Sparse Attention: Efficient Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/native-sparse-attention-efficient-long-context-llms/
30. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
31. AI Post Transformers: AWQ: On-Device LLM Compression and Acceleration — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/awq-on-device-llm-compression-and-acceleration/
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof