AI Post Transformers

ForkKV for Multi-LoRA Agent Serving


Listen Later

This episode explores ForkKV, a systems paper on serving multiple LoRA-based agents from one base language model without duplicating massive KV caches for shared context. It explains why ordinary prefix caching breaks once different LoRA adapters change the activations, then walks through the paper’s core idea: split cache state into a large shared base and a small adapter-specific residual, using an operating-system-style copy-on-write model for agent branches. The discussion connects that design to prior work on LoRA, prefix caching, PagedAttention, and disaggregated memory, making the argument that the real win is practical GPU memory efficiency for coding assistants and tool-using agent workflows. Listeners would find it interesting because it frames transformer serving as a memory-management problem and shows how borrowing ideas from Unix process forking could make multi-agent LLM systems far more scalable.
Sources:
1. ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache — Shao Wang, Rui Ren, Lin Gui, 2026
http://arxiv.org/abs/2604.06370
2. The UNIX Time-Sharing System — Dennis M. Ritchie and Ken Thompson, 1974
https://scholar.google.com/scholar?q=The+UNIX+Time-Sharing+System
3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
4. vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention — Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar, 2024
https://scholar.google.com/scholar?q=vAttention%3A+Dynamic+Memory+Management+for+Serving+LLMs+without+PagedAttention
5. ForkKV: Scaling Multi-LoRA Agent Serving via Copy-on-Write Disaggregated KV Cache — Shao Wang, Rui Ren, and Lin Gui, 2026
https://scholar.google.com/scholar?q=ForkKV%3A+Scaling+Multi-LoRA+Agent+Serving+via+Copy-on-Write+Disaggregated+KV+Cache
6. Efficiently Programming Large Language Models using SGLang — Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Jeff Huang, Chuyue Sun, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng, 2023
https://scholar.google.com/scholar?q=Efficiently+Programming+Large+Language+Models+using+SGLang
7. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
8. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan, 2024
https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool
9. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu, 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
10. LRAgent: efficient kv cache sharing for multi-lora llm agents — H. Jeon, H. Ha, and J. Kim, 2026
https://scholar.google.com/scholar?q=LRAgent%3A+efficient+kv+cache+sharing+for+multi-lora+llm+agents
11. Punica: multi-tenant lora serving — L. Chen, Z. Ye, Y. Wu, D. Zhuo, L. Ceze, and A. Krishnamurthy, 2024
https://scholar.google.com/scholar?q=Punica%3A+multi-tenant+lora+serving
12. S-LoRA: serving thousands of concurrent lora adapters — Y. Sheng, S. Cao, D. Li, C. Hooper, N. Lee, S. Yang, C. Chou, B. Zhu, L. Zheng, K. Keutzer, J. E. Gonzalez, and I. Stoica, 2023
https://scholar.google.com/scholar?q=S-LoRA%3A+serving+thousands+of+concurrent+lora+adapters
13. DLoRA: dynamically orchestrating requests and adapters for LoRA LLM serving — B. Wu, R. Zhu, Z. Zhang, P. Sun, X. Liu, and X. Jin, 2024
https://scholar.google.com/scholar?q=DLoRA%3A+dynamically+orchestrating+requests+and+adapters+for+LoRA+LLM+serving
14. Tokencake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications — Z. Bian, F. Wu, T. Ma, and Y. Zhuo, 2025
https://scholar.google.com/scholar?q=Tokencake%3A+A+KV-Cache-centric+Serving+Framework+for+LLM-based+Multi-Agent+Applications
15. KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent Workflows — Z. Pan, A. Patel, Z. Hu, Y. Shen, Y. Guan, W. Li, L. Qin, Y. Wang, and Y. Ding, 2025
https://scholar.google.com/scholar?q=KVFlow%3A+Efficient+Prefix+Caching+for+Accelerating+LLM-Based+Multi-Agent+Workflows
16. MobiLoRA: Accelerating LoRA-based LLM Inference on Mobile Devices via Context-aware KV Cache Optimization — Borui Li et al., 2025
https://scholar.google.com/scholar?q=MobiLoRA%3A+Accelerating+LoRA-based+LLM+Inference+on+Mobile+Devices+via+Context-aware+KV+Cache+Optimization
17. Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA — Allison Li, Kristjan Greenewald, Thomas Parnell, Navid Azizan, 2025
https://scholar.google.com/scholar?q=Efficient+Multi-Adapter+LLM+Serving+via+Cross-Model+KV-Cache+Reuse+with+Activated+LoRA
18. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
19. AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization — Qiyang Li et al., 2026
https://scholar.google.com/scholar?q=AdaFuse%3A+Accelerating+Dynamic+Adapter+Inference+via+Token-Level+Pre-Gating+and+Fused+Kernel+Optimization
20. ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference — Xiang Liu et al., 2025
https://scholar.google.com/scholar?q=ChunkKV%3A+Semantic-Preserving+KV+Cache+Compression+for+Efficient+Long-Context+LLM+Inference
21. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2025
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
22. ELORA: Efficient LoRA and KV Cache Management for Multi-LoRA LLM Serving — Jiuchen Shi et al., 2026
https://scholar.google.com/scholar?q=ELORA%3A+Efficient+LoRA+and+KV+Cache+Management+for+Multi-LoRA+LLM+Serving
23. AI Post Transformers: Efficient KV Cache Sharing for Multi-LoRA Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-efficient-kv-cache-sharing-for-multi-lor-afda05.mp3
24. AI Post Transformers: TokenDance for Multi-Agent KV Cache Sharing — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tokendance-for-multi-agent-kv-cache-shar-aa9b99.mp3
25. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
26. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
27. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
Interactive Visualization: ForkKV for Multi-LoRA Agent Serving
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof