This episode explores a 2026 paper proposing that multiple LLM agents should share transformer KV-cache state, not just text, so they can avoid repeatedly paying the prefill cost of rereading the same plans, critiques, and intermediate outputs. It explains the systems background behind prefix caching, vLLM’s PagedAttention, and SGLang, then focuses on why multi-agent workflows break the exact-prefix assumption and make segment-level reuse much harder. The discussion highlights the paper’s core technical tension: the idea is compelling, but reusing cached activations across different prompt positions is fragile because of positional encoding effects such as RoPE misalignment and attention behavior. Listeners would find it interesting because it connects a practical bottleneck in agent systems to deep transformer internals, while also questioning whether the paper truly delivers fine-grained semantic sharing or a narrower form of reusable output caching.
Sources:
1. Breaking the Prefix Barrier with Shared KV Cache
https://openreview.net/forum?id=kgzBkyqg6Z
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, Ying Sheng, 2024
https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs
4. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2025
https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion
5. KVCOMM: Online Cross-context KV-cache Communication for Efficient LLM-based Multi-agent Systems — Hancheng Ye, Zhengqi Gao, Mingyuan Ma, Qinsi Wang, Yuzhe Fu, Ming-Yu Chung, Yueqian Lin, Zhijian Liu, Jianyi Zhang, Danyang Zhuo, Yiran Chen, 2025
https://scholar.google.com/scholar?q=KVCOMM%3A+Online+Cross-context+KV-cache+Communication+for+Efficient+LLM-based+Multi-agent+Systems
6. EPIC: Efficient Position-Independent Caching for Serving Large Language Models — Junhao Hu, Wenrui Huang, Weidong Wang, Haoyi Wang, Tiancheng Hu, Qin Zhang, Hao Feng, Xusheng Chen, Yizhou Shan, Tao Xie, 2025
https://scholar.google.com/scholar?q=EPIC%3A+Efficient+Position-Independent+Caching+for+Serving+Large+Language+Models
7. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan, Ajjkumar Patel, Zhengding Hu, Yipeng Shen, Yue Guan, Wan-Lu Li, Lianhui Qin, Yida Wang, Yufei Ding, 2025
https://scholar.google.com/scholar?q=KVFlow%3A+Efficient+Prefix+Caching+for+Accelerating+LLM-Based+Multi-Agent+Workflows
8. DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving — Yuhan Liu, Yuyang Huang, Jiayi Yao, Shaoting Feng, Zhuohan Gu, Kuntai Du, Hanchen Li, Yihua Cheng, Junchen Jiang, Shan Lu, Madan Musuvathi, Esha Choukse, 2024
https://scholar.google.com/scholar?q=DroidSpeak%3A+KV+Cache+Sharing+for+Cross-LLM+Communication+and+Multi-LLM+Serving
9. TokenDance: Scaling Multi-Agent LLM Serving via Collective KV Cache Sharing — Zhuohang Bian, Feiyang Wu, Chengrui Zhang, Hangcheng Dong, Yun Liang, Youwei Zhuo, 2026
https://scholar.google.com/scholar?q=TokenDance%3A+Scaling+Multi-Agent+LLM+Serving+via+Collective+KV+Cache+Sharing
10. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — authors unclear from Scholar snippet, 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
11. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — authors unclear from Scholar snippet, 2025
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
12. Cache-craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation — authors unclear from Scholar snippet, 2025
https://scholar.google.com/scholar?q=Cache-craft%3A+Managing+Chunk-Caches+for+Efficient+Retrieval-Augmented+Generation
13. AttentionStore: Cost-Effective Attention Reuse Across Multi-Turn Conversations in Large Language Model Serving — authors unclear from Scholar snippet, 2025
https://scholar.google.com/scholar?q=AttentionStore%3A+Cost-Effective+Attention+Reuse+Across+Multi-Turn+Conversations+in+Large+Language+Model+Serving
14. BAT: Efficient Generative Recommender Serving with Bipartite Attention — authors unclear from Scholar snippet, 2025
https://scholar.google.com/scholar?q=BAT%3A+Efficient+Generative+Recommender+Serving+with+Bipartite+Attention
15. AI Post Transformers: Efficient KV Cache Sharing for Multi-LoRA Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-efficient-kv-cache-sharing-for-multi-lor-afda05.mp3
16. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3
17. AI Post Transformers: KV Cache TTL for Multi-Turn Agent Scheduling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-kv-cache-ttl-for-multi-turn-agent-schedu-996bf1.mp3
18. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
19. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
Interactive Visualization: Breaking the Prefix Barrier with Shared KV Cache