This episode explores the paper Cartridges at Scale, which asks whether large document collections can be distilled into reusable modular KV-cache memories so a model can answer questions without repeatedly rereading raw text. It explains what a cartridge is, how context distillation turns full-document context into compact learned prefixes, and why that differs from prompt caching, fine-tuning, ordinary long-context prompting, and text RAG. The discussion centers on the paper’s main claim that per-document memories do not reliably compose when trained independently, so the authors jointly train cartridges with both relevant and irrelevant memories present to teach a frozen model which compressed document to attend to in a noisy multi-document setting. Listeners would find it interesting because it treats the KV cache as a potential external memory layer that could reduce inference cost and latency while exposing hard questions about compositionality, transparency, and whether learned memory modules can outperform standard retrieval pipelines.
Sources:
1. Cartridges at Scale: Training Modular KV Caches over Large Document Collections — Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert, 2026
http://arxiv.org/abs/2606.04557
2. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Douwe Kiela, et al., 2020
https://arxiv.org/abs/2005.11401
3. Prompt Cache: Modular Attention Reuse for Low-Latency Inference — In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, Lin Zhong, 2023
https://arxiv.org/abs/2311.04934
4. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study — Sabri Eyuboglu, Ryan Ehrlich, Simran Arora, Neel Guha, James Zou, Azalia Mirhoseini, Christopher Re, et al., 2025
https://arxiv.org/abs/2506.06266
5. Cartridges at Scale: Training Modular KV Caches over Large Document Collections — Momchil Hardalov, Gonzalo Iglesias, Adrià de Gispert, 2026
https://arxiv.org/abs/2606.04557
6. Cartridges: Lightweight and General-Purpose Long Context Representations via Self-Study (https://arxiv.org/abs/2506.06266) — Sabri Eyuboglu, Ryan S. Ehrlich, Simran Arora, Neel Guha, Dylan Zinsley, Emily R. Liu, Atri Rudra, James Y. Zou, Azalia Mirhoseini, Christopher Re, 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+General-Purpose+Long+Context+Representations+via+Self-Study+%28https%3A%2F%2Farxiv.org%2Fabs%2F2506.06266%29
7. Learned Structure in CARTRIDGES: Keys as Shareable Routers in Self-Studied Representations (https://arxiv.org/abs/2508.17032) — Maurizio Diaz, 2025
https://scholar.google.com/scholar?q=Learned+Structure+in+CARTRIDGES%3A+Keys+as+Shareable+Routers+in+Self-Studied+Representations+%28https%3A%2F%2Farxiv.org%2Fabs%2F2508.17032%29
8. xRAG: Extreme Context Compression for Retrieval-Augmented Generation with One Token (https://arxiv.org/abs/2405.13792) — Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, Dongyan Zhao, 2024
https://scholar.google.com/scholar?q=xRAG%3A+Extreme+Context+Compression+for+Retrieval-Augmented+Generation+with+One+Token+%28https%3A%2F%2Farxiv.org%2Fabs%2F2405.13792%29
9. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction (https://arxiv.org/abs/2505.23416) — Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee, Sangdoo Yun, Hyun Oh Song, 2025
https://scholar.google.com/scholar?q=KVzip%3A+Query-Agnostic+KV+Cache+Compression+with+Context+Reconstruction+%28https%3A%2F%2Farxiv.org%2Fabs%2F2505.23416%29
10. T2-RAGBench: Text-and-Table Benchmark for Evaluating Retrieval-Augmented Generation (https://aclanthology.org/2026.eacl-long.8/) — Jan Strich, Enes Kutay Isgorur, Maximilian Trescher, Chris Biemann, Martin Semmann, 2026
https://scholar.google.com/scholar?q=T2-RAGBench%3A+Text-and-Table+Benchmark+for+Evaluating+Retrieval-Augmented+Generation+%28https%3A%2F%2Faclanthology.org%2F2026.eacl-long.8%2F%29
11. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
12. Hierarchical Document Refinement for Long-context Retrieval-augmented Generation — Jiajie Jin et al., 2025
https://scholar.google.com/scholar?q=Hierarchical+Document+Refinement+for+Long-context+Retrieval-augmented+Generation
13. LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing — Kuan Li et al., 2025
https://scholar.google.com/scholar?q=LaRA%3A+Benchmarking+Retrieval-Augmented+Generation+and+Long-Context+LLMs+--+No+Silver+Bullet+for+LC+or+RAG+Routing
14. ChatQA 2: Bridging the Gap to Proprietary LLMs in Long Context and RAG Capabilities — Peng Xu et al., 2024
https://scholar.google.com/scholar?q=ChatQA+2%3A+Bridging+the+Gap+to+Proprietary+LLMs+in+Long+Context+and+RAG+Capabilities
15. LegalBench-RAG: A Benchmark for Retrieval-Augmented Generation in the Legal Domain — Nicholas Pipitone and Ghita Houir Alami, 2024
https://scholar.google.com/scholar?q=LegalBench-RAG%3A+A+Benchmark+for+Retrieval-Augmented+Generation+in+the+Legal+Domain
16. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025
https://scholar.google.com/scholar?q=KVShare%3A+An+LLM+Service+System+with+Efficient+and+Effective+Multi-Tenant+KV+Cache+Reuse
17. AI Post Transformers: KVzip for Query-Agnostic KV Cache Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-29-kvzip-for-query-agnostic-kv-cache-compre-72afe5.mp3
18. AI Post Transformers: From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-22-from-prefix-cache-to-fusion-rag-9c5d39.mp3
19. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3
20. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3
21. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
22. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3