This episode explores Information-Aware KV Cache Compression for Long Reasoning, a paper about making long-context inference cheaper and more reliable by deciding which KV-cache tokens to keep during extended reasoning. It explains why long prefilling and long decoding turn the cache into a major memory bottleneck, and why common heuristics such as sliding windows or recent-attention-based retention can discard tokens that only become important much later. The discussion centers on the paper’s claim that future usefulness is better captured by information-theoretic signals like predictive entropy and Forward Influence, with experiments showing that attention-ranked tokens help short-horizon predictions while entropy-ranked tokens matter more over long horizons. Listeners get a concrete account of how InfoKV blends recent attention with per-layer entropy-based scoring to improve the tradeoff between memory savings and long-range reasoning quality.
Sources:
1. Information-Aware KV Cache Compression for Long Reasoning — Jushi Kai, Zhuiri Xiao, Alexandra Birch, Zhouhan Lin, 2026
http://arxiv.org/abs/2606.26875
2. Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time — Zichang Liu, Aditya Desai, Fangshuo Liao, Anshumali Shrivastava, et al., 2023
https://scholar.google.com/scholar?q=Scissorhands%3A+Exploiting+the+Persistence+of+Importance+Hypothesis+for+LLM+KV+Cache+Compression+at+Test+Time
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Beidi Chen, Christopher Re, et al., 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. SnapKV: LLM Knows What You are Looking for Before Generation — Yuhong Li, Yingbing Huang, Bowen Yang, Patrick Lewis, Deming Chen, et al., 2024
https://scholar.google.com/scholar?q=SnapKV%3A+LLM+Knows+What+You+are+Looking+for+Before+Generation
5. Information-Aware KV Cache Compression for Long Reasoning — Jushi Kai, Zhuiri Xiao, Alexandra Birch, Zhouhan Lin, 2026
https://scholar.google.com/scholar?q=Information-Aware+KV+Cache+Compression+for+Long+Reasoning
6. Expected Attention: KV Cache Compression by Estimating Attention from Future Queries Distribution — Alessio Devoto, Maximilian Jeblick, Simon Jegou, 2025
https://scholar.google.com/scholar?q=Expected+Attention%3A+KV+Cache+Compression+by+Estimating+Attention+from+Future+Queries+Distribution
7. Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning — Jiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon Kim, 2025
https://scholar.google.com/scholar?q=Reasoning+Path+Compression%3A+Compressing+Generation+Trajectories+for+Efficient+LLM+Reasoning
8. Compressing Context to Enhance Inference Efficiency of Large Language Models — Yucheng Li, Bo Dong, Chenghua Lin, Frank Guerin, 2023
https://scholar.google.com/scholar?q=Compressing+Context+to+Enhance+Inference+Efficiency+of+Large+Language+Models
9. FreqKV: Key-Value Compression in Frequency Domain for Context Window Extension — Jushi Kai et al., 2026
https://scholar.google.com/scholar?q=FreqKV%3A+Key-Value+Compression+in+Frequency+Domain+for+Context+Window+Extension
10. LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion — Zhan Ling et al., 2025
https://scholar.google.com/scholar?q=LongReason%3A+A+Synthetic+Long-Context+Reasoning+Benchmark+via+Context+Expansion
11. Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval — Yuwei Zhang et al., 2025
https://scholar.google.com/scholar?q=Attention+Reveals+More+Than+Tokens%3A+Training-Free+Long-Context+Reasoning+with+Attention-guided+Retrieval
12. Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions — Sungmin Kang et al., 2025
https://scholar.google.com/scholar?q=Uncertainty+Quantification+for+Hallucination+Detection+in+Large+Language+Models%3A+Foundations%2C+Methodology%2C+and+Future+Directions
13. Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations — Christian Tomani et al., 2024
https://scholar.google.com/scholar?q=Uncertainty-Based+Abstention+in+LLMs+Improves+Safety+and+Reduces+Hallucinations
14. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction
15. Can LLMs Maintain Fundamental Abilities under KV Cache Compression? — Xiang Liu et al., 2025
https://scholar.google.com/scholar?q=Can+LLMs+Maintain+Fundamental+Abilities+under+KV+Cache+Compression%3F
16. KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches — Jiayi Yuan et al., 2024
https://scholar.google.com/scholar?q=KV+Cache+Compression%2C+But+What+Must+We+Give+in+Return%3F+A+Comprehensive+Benchmark+of+Long+Context+Capable+Approaches
17. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
18. AI Post Transformers: IndexMem: Learned KV-Cache Eviction for Long-Context LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-12-indexmem-learned-kv-cache-eviction-for-l-132c2a.mp3
19. AI Post Transformers: When Quantization Hurts Reasoning Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-17-when-quantization-hurts-reasoning-models-eca9e7.mp3
20. AI Post Transformers: Hyper-Scaling LLM Inference with KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/hyper-scaling-llm-inference-with-kv-cache-compression/
21. AI Post Transformers: Lattice: Fixed-Slot Compression for Transformer Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-lattice-fixed-slot-compression-for-trans-5509ea.mp3
22. AI Post Transformers: Adaptive Compression Techniques for Efficient LLM Inference — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/adaptive-compression-techniques-for-efficient-llm-inference/
23. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
24. AI Post Transformers: When LoRA Helps Under KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-12-when-lora-helps-under-kv-cache-compressi-76dda6.mp3