This episode explores CacheFlow, a systems approach to speeding up long-context LLM serving by restoring transformer KV caches more intelligently. It explains the tradeoffs between recomputing prior attention state, loading it from storage, or combining both, and argues that the real user-facing bottleneck is now time-to-first-token rather than raw generation speed. The discussion focuses on CacheFlow’s main idea: a batch-aware scheduler that splits restoration across recomputation and I/O at token, layer, and GPU levels to reduce wasted work under contention. Listeners would find it interesting because it shows how practical transformer serving is increasingly shaped by runtime scheduling, cache movement, and latency engineering rather than new model architectures.
Sources:
1. CacheFlow and 3D-Parallel KV Cache Restoration
https://arxiv.org/pdf/2604.25080
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention — Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, Pengfei Zuo, 2024
https://scholar.google.com/scholar?q=Cost-Efficient+Large+Language+Model+Serving+for+Multi-turn+Conversations+with+CachedAttention
4. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
5. Fast State Restoration in LLM Serving with HCache — Shiwei Gao, Youmin Chen, Jiwu Shu, 2024
https://scholar.google.com/scholar?q=Fast+State+Restoration+in+LLM+Serving+with+HCache
6. LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference — Yihua Cheng, Yuhan Liu, Jiayi Yao, Yuwei An, Xiaokun Chen, Shaoting Feng, Yuyang Huang, Samuel Shen, Kuntai Du, Junchen Jiang, 2025
https://scholar.google.com/scholar?q=LMCache%3A+An+Efficient+KV+Cache+Layer+for+Enterprise-Scale+LLM+Inference
7. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
8. Mooncake: Trading More Storage for Less Computation — A KVCache-centric Architecture for Serving LLM Chatbot — Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2025
https://scholar.google.com/scholar?q=Mooncake%3A+Trading+More+Storage+for+Less+Computation+%E2%80%94+A+KVCache-centric+Architecture+for+Serving+LLM+Chatbot
9. DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving — Foteini Strati, Sara Mcallister, Amar Phanishayee, Jakub Tarnawski, Ana Klimovic, 2024
https://scholar.google.com/scholar?q=D%C3%A9j%C3%A0Vu%3A+KV-cache+Streaming+for+Fast%2C+Fault-tolerant+Generative+LLM+Serving
10. KV Prediction for Improved Time to First Token — Maxwell Horton, Qingqing Cao, Chenfan Sun, Yanzi Jin, Sachin Mehta, Mohammad Rastegari, Moin Nabi, 2025
https://scholar.google.com/scholar?q=KV+Prediction+for+Improved+Time+to+First+Token
11. KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing — Yifei Yang et al., 2024
https://scholar.google.com/scholar?q=KVSharer%3A+Efficient+Inference+via+Layer-Wise+Dissimilar+KV+Cache+Sharing
12. CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing — Yixuan Wang et al., 2025
https://scholar.google.com/scholar?q=CommonKV%3A+Compressing+KV+Cache+with+Cross-layer+Parameter+Sharing
13. Lossless KV Cache Compression to 2% — Zhen Yang et al., 2024
https://scholar.google.com/scholar?q=Lossless+KV+Cache+Compression+to+2%25
14. XKV: Personalized KV Cache Memory Reduction for Long-Context LLM Inference — Weizhuo Li et al., 2024
https://scholar.google.com/scholar?q=XKV%3A+Personalized+KV+Cache+Memory+Reduction+for+Long-Context+LLM+Inference
15. TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding — Hanshi Sun et al., 2024
https://scholar.google.com/scholar?q=TriForce%3A+Lossless+Acceleration+of+Long+Sequence+Generation+with+Hierarchical+Speculative+Decoding
16. MagicDec: Breaking the Latency-Throughput Tradeoff for Long Context Generation with Speculative Decoding — Jian Chen, Vashisth Tiwari, Ranajoy Sadhukhan, Zhuoming Chen, Jinyuan Shi, Ian En-Hsu Yen, Beidi Chen, 2024
https://scholar.google.com/scholar?q=MagicDec%3A+Breaking+the+Latency-Throughput+Tradeoff+for+Long+Context+Generation+with+Speculative+Decoding
17. SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs — Shibo Jie et al., 2025
https://scholar.google.com/scholar?q=SpeCache%3A+Speculative+Key-Value+Caching+for+Efficient+Generation+of+LLMs
18. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving — Chao Wang, Pengfei Zuo, Zhangyu Chen, Yunkai Liang, Zhou Yu, Ming-Chang Yang, 2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving
19. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
20. AI Post Transformers: ScoutAttention for Efficient KV Cache Offloading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-scoutattention-for-efficient-kv-cache-of-b26699.mp3
21. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
22. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
23. AI Post Transformers: KV Cache TTL for Multi-Turn Agent Scheduling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-kv-cache-ttl-for-multi-turn-agent-schedu-996bf1.mp3
24. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
25. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
Interactive Visualization: CacheFlow and 3D-Parallel KV Cache Restoration