AI Post Transformers

Training Modular KV Caches at Scale


Listen Later

This episode explores the paper Cartridges at Scale, which asks whether large document collections can be distilled into reusable modular KV-cache memories so a model can answer questions without repeatedly rereading raw text. It explains what a cartridge is, how context distillation turns full-document context into compact learned prefixes, and why that differs from prompt caching, fine-tuning, ordinary long-context prompting, and text RAG. The discussion centers on the paper’s main claim that per-document memories do not reliably compose when trained independently, so the authors jointly train cartridges with both relevant and irrelevant memories present to teach a frozen model which compressed document to attend to in a noisy multi-document setting. Listeners would find it interesting because it treats the KV cache as a potential external memory layer that could reduce inference cost and latency while exposing hard questions about compositionality, transparency, and whether learned memory modules can outperform standard retrieval pipelines.
Sources:
1. Cartridges at Scale: Training Modular KV Caches over Large Document Collections — Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert, 2026
http://arxiv.org/abs/2606.04557
2. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Douwe Kiela, et al., 2020
https://arxiv.org/abs/2005.11401
3. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong, 2023
https://arxiv.org/abs/2311.04934
4. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, James Zou, Azalia Mirhoseini, Christopher Re, et al., 2025
https://arxiv.org/abs/2506.06266
5. Cartridges at Scale: Training Modular KV Caches over Large Document Collections — Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert, 2026
https://arxiv.org/abs/2606.04557
6. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study (https://arxiv.org/abs/2506.06266) — Sabri Eyuboglu, Ryan S. Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily R. Liu, Atri Rudra, James Y. Zou, Azalia Mirhoseini, Christopher Re, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study+%28https%3A%2F%2Farxiv.org%2Fabs%2F2506.06266%29
7. Learned Structure in CARTRIDGES: Keys as Shareable Routers in Self-Studied Representations (https://arxiv.org/abs/2508.17032) — Maurizio Diaz, 2025
https://scholar.google.com/scholar?q=Learned+Structure+in+CARTRIDGES%3A+Keys+as+Shareable+Routers+in+Self-Studied+Representations+%28https%3A%2F%2Farxiv.org%2Fabs%2F2508.17032%29
8. xRAG: Extreme Context Compression for Retrieval-Augmented Generation with One Token (https://arxiv.org/abs/2405.13792) — Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, Dongyan Zhao, 2024
https://scholar.google.com/scholar?q=xRAG%3A+Extreme+Context+Compression+for+Retrieval-Augmented+Generation+with+One+Token+%28https%3A%2F%2Farxiv.org%2Fabs%2F2405.13792%29
9. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction (https://arxiv.org/abs/2505.23416) — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction+%28https%3A%2F%2Farxiv.org%2Fabs%2F2505.23416%29
10. T2-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation (https://aclanthology.org/2026.eacl-long.8/) — Jan Strich, Enes Kutay Isgorur, Maximilian Trescher, Chris Biemann, Martin Semmann, 2026
https://scholar.google.com/scholar?q=T2-RAGBench%3A+Text-and-Table+Benchmark+for+Evaluating+Retrieval-Augmented+Generation+%28https%3A%2F%2Faclanthology.org%2F2026.eacl-long.8%2F%29
11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
12. Hierarchical Document Refinement for Long-context Retrieval-augmented Generation — Jiajie Jin et al., 2025
https://scholar.google.com/scholar?q=Hierarchical+Document+Refinement+for+Long-context+Retrieval-augmented+Generation
13. LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing — Kuan Li et al., 2025
https://scholar.google.com/scholar?q=LaRA%3A+Benchmarking+Retrieval-Augmented+Generation+and+Long-Context+LLMs+--+No+Silver+Bullet+for+LC+or+RAG+Routing
14. ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities — Peng Xu et al., 2024
https://scholar.google.com/scholar?q=ChatQA+2%3A+Bridging+the+Gap+to+Proprietary+LLMs+in+Long+Context+and+RAG+Capabilities
15. LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain — Nicholas Pipitone and Ghita Houir Alami, 2024
https://scholar.google.com/scholar?q=LegalBench-RAG%3A+A+Benchmark+for+Retrieval-Augmented+Generation+in+the+Legal+Domain
16. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025
https://scholar.google.com/scholar?q=KVShare%3A+An+LLM+Service+System+with+Efficient+and+Effective+Multi-Tenant+KV+Cache+Reuse
17. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3
18. AI Post Transformers: From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-22-from-prefix-cache-to-fusion-rag-9c5d39.mp3
19. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3
20. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3
21. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
22. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof