This episode explores Mooncake, a production LLM serving architecture that treats KV cache reuse and movement as the central challenge in long-context chat, not just raw GPU compute. It explains why prefill and decode stress hardware in different ways, how metrics like time to first token and time between tokens drive system design, and why separating those phases helps meet real latency targets. The discussion walks through Mooncake’s cache-first scheduler, tiered KV storage across GPU memory, CPU DRAM, and SSD, and its use of RDMA, chunked prefill, and layer-wise overlap to start decoding sooner while reusing existing state. It also argues that Mooncake’s interest lies less in a single breakthrough than in how it combines prefix-aware routing, overload-aware early rejection, and cross-node KV reuse into a practical serving stack for large-scale chat systems.
Sources:
1. Mooncake for KV Cache-Centric LLM Serving
https://arxiv.org/pdf/2407.00079
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. SGLang: Efficient Execution of Structured Language Model Programs — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Ying Sheng, et al., 2023
https://scholar.google.com/scholar?q=SGLang%3A+Efficient+Execution+of+Structured+Language+Model+Programs
4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
6. Overload Control for Scaling WeChat Microservices — Hao Zhou, Ming Chen, Qian Lin, Yong Wang, Xiaobin She, Sifan Liu, Rui Gu, Beng Chin Ooi, Junfeng Yang, 2018
https://scholar.google.com/scholar?q=Overload+Control+for+Scaling+WeChat+Microservices
7. Overload Control for microsecond-scale RPCs with Breakwater — Inho Cho, Ahmed Saeed, Joshua Fried, Seo Jin Park, Mohammad Alizadeh, Adam Belay, 2020
https://scholar.google.com/scholar?q=Overload+Control+for+microsecond-scale+RPCs+with+Breakwater
8. SCORPIO: Serving the Right Requests at the Right Time for Heterogeneous SLOs in LLM Inference — Yinghao Tang, Tingfeng Lan, Xiuqi Huang, Hui Lu, Wei Chen, 2025
https://scholar.google.com/scholar?q=SCORPIO%3A+Serving+the+Right+Requests+at+the+Right+Time+for+Heterogeneous+SLOs+in+LLM+Inference
9. Splitwise: Efficient Generative LLM Inference Using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Inigo Goiri, Aashaka Shah, Saeed Maleki, Ricardo Bianchini, 2023
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+Using+Phase+Splitting
10. AttentionStore: Cost-Effective Attention Reuse Across Multi-Turn Conversations in Large Language Model Serving — Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, Pengfei Zuo, 2024
https://scholar.google.com/scholar?q=AttentionStore%3A+Cost-Effective+Attention+Reuse+Across+Multi-Turn+Conversations+in+Large+Language+Model+Serving
11. Preble: Efficient Distributed Prompt Scheduling for LLM Serving — Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, Yiying Zhang, 2024
https://scholar.google.com/scholar?q=Preble%3A+Efficient+Distributed+Prompt+Scheduling+for+LLM+Serving
12. CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion — Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, Junchen Jiang, 2024
https://scholar.google.com/scholar?q=CacheBlend%3A+Fast+Large+Language+Model+Serving+for+RAG+with+Cached+Knowledge+Fusion
13. P/D-Serve: Serving Disaggregated Large Language Model at Scale — Yibo Jin et al., 2024
https://scholar.google.com/scholar?q=P%2FD-Serve%3A+Serving+Disaggregated+Large+Language+Model+at+Scale
14. POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference — Aditya K. Kamath et al., 2025
https://scholar.google.com/scholar?q=POD-Attention%3A+Unlocking+Full+Prefill-Decode+Overlap+for+Faster+LLM+Inference
15. Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models — Siyan Zhao et al., 2024
https://scholar.google.com/scholar?q=Prepacking%3A+A+Simple+Method+for+Fast+Prefilling+and+Increased+Throughput+in+Large+Language+Models
16. Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference — Donghyeon Joo et al., 2025
https://scholar.google.com/scholar?q=Mustafar%3A+Promoting+Unstructured+Sparsity+for+KV+Cache+Pruning+in+LLM+Inference
17. SparK: Query-Aware Unstructured Sparsity with Recoverable KV Cache Channel Pruning — Huanxuan Liao et al., 2025/2026
https://scholar.google.com/scholar?q=SparK%3A+Query-Aware+Unstructured+Sparsity+with+Recoverable+KV+Cache+Channel+Pruning
18. More Tokens, Lower Precision: Towards the Optimal Token-Precision Trade-off in KV Cache Compression — Jiebin Zhang et al., 2024
https://scholar.google.com/scholar?q=More+Tokens%2C+Lower+Precision%3A+Towards+the+Optimal+Token-Precision+Trade-off+in+KV+Cache+Compression
19. KVShare: An LLM Service System with Efficient and Effective Multi-Tenant KV Cache Reuse — Huan Yang et al., 2025
https://scholar.google.com/scholar?q=KVShare%3A+An+LLM+Service+System+with+Efficient+and+Effective+Multi-Tenant+KV+Cache+Reuse
20. CacheSolidarity: Preventing Prefix Caching Side Channels in Multi-tenant LLM Serving Systems — Panagiotis Georgios Pennas et al., 2026
https://scholar.google.com/scholar?q=CacheSolidarity%3A+Preventing+Prefix+Caching+Side+Channels+in+Multi-tenant+LLM+Serving+Systems
21. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
22. AI Post Transformers: Characterizing LLM KV Cache Workloads in Production — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/characterizing-llm-kv-cache-workloads-in-production/
23. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
24. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3
25. AI Post Transformers: Beluga: CXL Memory Pooling for LLM KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-27-beluga-cxl-memory-pooling-for-llm-kv-cac-b6142f.mp3
26. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
Interactive Visualization: Mooncake for KV Cache-Centric LLM Serving