This episode explores Harvest, a system for LLM inference that uses idle HBM on neighboring NVLink-connected GPUs as a temporary cache when a serving GPU runs out of local memory. It explains why LLM serving is often bottlenecked more by memory capacity and data movement than by raw compute, focusing on two concrete cases: growing KV caches during long-context decoding and the shifting expert weights used in mixture-of-experts models. A key argument is that peer GPU memory is only useful if it is revocable without breaking correctness, so Harvest treats borrowed memory as a best-effort cache backed by authoritative copies or reconstruction paths elsewhere. Listeners get specific performance results, including up to 5.65x lower KV-cache transfer latency than CPU offload, 7.5x to 9.5x faster expert transfers over NVLink, and roughly 1.5x to 2.0x throughput gains on models such as Qwen and Phi-3.5.
Sources:
1. Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference — Nikhil Gopal, Kostis Kaffes, 2026
http://arxiv.org/abs/2602.00328
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU — Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Beidi Chen, Percy Liang, Christopher Re, Ion Stoica, Ce Zhang, 2023
https://scholar.google.com/scholar?q=FlexGen%3A+High-Throughput+Generative+Inference+of+Large+Language+Models+with+a+Single+GPU
4. Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models — Keisuke Kamahori, Tian Tang, Yile Gu, Kan Zhu, Baris Kasikci, 2025
https://scholar.google.com/scholar?q=Fiddler%3A+CPU-GPU+Orchestration+for+Fast+Inference+of+Mixture-of-Experts+Models
5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu, et al., 2025
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
6. AQUA: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU Domains — Abhishek Vijaya Kumar, Gianni Antichi, Rachee Singh, 2025
https://scholar.google.com/scholar?q=AQUA%3A+Network-Accelerated+Memory+Offloading+for+LLMs+in+Scale-Up+GPU+Domains
7. Accurate Expert Predictions in MoE Inference via Cross-Layer Gate — Zhiyuan Fang, Hong Huang, Yiming Lyu, Jiyang Chen, Yu Yu, Zexi Zheng, 2025
https://scholar.google.com/scholar?q=Accurate+Expert+Predictions+in+MoE+Inference+via+Cross-Layer+Gate
8. Characterization of Large Language Model Development in the Datacenter — Qinghao Hu et al., 2024
https://scholar.google.com/scholar?q=Characterization+of+Large+Language+Model+Development+in+the+Datacenter
9. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters — Qizhen Weng et al., 2022
https://scholar.google.com/scholar?q=MLaaS+in+the+Wild%3A+Workload+Analysis+and+Scheduling+in+Large-Scale+Heterogeneous+GPU+Clusters
10. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu et al., 2024
https://arxiv.org/abs/2406.17565
11. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Zaifeng Pan et al., 2025
https://arxiv.org/abs/2507.07400
12. Learned Prefix Caching for Efficient LLM Inference — Dongsheng Yang et al., 2025
https://papers.neurips.cc/paper_files/paper/2025/hash/414f642a1ea9350006669774cba9bcd4-Abstract-Conference.html
13. DuoServe-MoE: Dual-Phase Expert Prefetch and Caching for LLM Inference QoS Assurance — Yuning Zhang et al., 2025
https://arxiv.org/abs/2509.07379
14. ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference — Zixu Shen et al., 2025
https://arxiv.org/abs/2510.26730
15. AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference — Shuzhang Zhong et al., 2024
https://arxiv.org/abs/2408.10284
16. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023
https://arxiv.org/abs/2310.05869
17. Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent Memory — Aydar Bulatov et al., 2024
https://ojs.aaai.org/index.php/AAAI/article/download/29722/31239
18. AI Post Transformers: ScoutAttention for Efficient KV Cache Offloading — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-scoutattention-for-efficient-kv-cache-of-b26699.mp3
19. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
20. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
21. AI Post Transformers: FlexGen: High-Throughput LLM Inference on a Single GPU — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/flexgen-high-throughput-llm-inference-on-a-single-gpu/
22. AI Post Transformers: Stochastic KV Routing for Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stochastic-kv-routing-for-cache-sharing-5fef63.mp3
23. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3