AI Post Transformers

ScoutAttention for Efficient KV Cache Offloading


Listen Later

This episode explores ScoutAttention, a systems paper on speeding up long-context LLM inference by managing KV cache growth more intelligently instead of treating it as a pure memory-capacity problem. It explains why long prompts, retrieval-heavy inputs, and extended reasoning traces make decoding increasingly constrained by memory traffic, and why simply offloading cache data to CPU memory can still leave GPUs stalled. The discussion focuses on the paper’s core idea: keep dense, high-speed attention on GPU, let the CPU handle only a pruned sparse set of offloaded KV blocks, and use layer-ahead CPU pre-computation plus asynchronous overlap to reduce waiting. Listeners would find it interesting because it frames transformer inference as a hardware scheduling problem and shows how throughput gains can come from smarter coordination between GPU compute, CPU compute, and data movement rather than from changing the model itself.
Sources:
1. ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference — Qiuyang Zhang, Kai Zhou, Ding Tang, Kai Lu, Cheng Li, Zhenyu Yang, Peng Xu, Jiguang Wan, 2026
http://arxiv.org/abs/2603.27138
2. InfiniGen — Not specified in the provided excerpt, Not specified in the provided excerpt
https://scholar.google.com/scholar?q=InfiniGen
3. HGCA — Not specified in the provided excerpt, Not specified in the provided excerpt
https://scholar.google.com/scholar?q=HGCA
4. OpenAI o1 — OpenAI, Not specified in the provided excerpt
https://scholar.google.com/scholar?q=OpenAI+o1
5. DeepSeek-R1 — DeepSeek, Not specified in the provided excerpt
https://scholar.google.com/scholar?q=DeepSeek-R1
6. Retrieval-Augmented Generation — Not specified in the provided excerpt, Not specified in the provided excerpt
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation
7. KV Cache Offloading for Context-Intensive Tasks — Andrey Bocharnikov, Ivan Ermakov, Denis Kuznedelev, Vyacheslav Zhdanovskiy, Yegor Yershov, 2026
https://scholar.google.com/scholar?q=KV+Cache+Offloading+for+Context-Intensive+Tasks
8. In-context KV-Cache Eviction for LLMs via Attention-Gate — Zihao Zeng, Bokai Lin, Tianqi Hou, Hao Zhang, Zhijie Deng, 2024
https://scholar.google.com/scholar?q=In-context+KV-Cache+Eviction+for+LLMs+via+Attention-Gate
9. LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation — Jinwoo Ahn, Ingyu Seong, Akhil Kedia, Junhan Kim, Hyemi Jang, Kangwook Lee, Yongkweon Jeon, 2026
https://scholar.google.com/scholar?q=LookaheadKV%3A+Fast+and+Accurate+KV+Cache+Eviction+by+Glimpsing+into+the+Future+without+Generation
10. MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse Attention — Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H. Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, Lili Qiu, 2024
https://scholar.google.com/scholar?q=MInference+1.0%3A+Accelerating+Pre-filling+for+Long-Context+LLMs+via+Dynamic+Sparse+Attention
11. SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention — Qianchao Zhu, Jiangfei Duan, Chang Chen, Siran Liu, Xiuhong Li, Guanyu Feng, Xin Lv, Huanqi Cao, Xiao Chuanfu, Xingcheng Zhang, Dahua Lin, Chao Yang, 2024
https://scholar.google.com/scholar?q=SampleAttention%3A+Near-Lossless+Acceleration+of+Long+Context+LLM+Inference+with+Adaptive+Structured+Sparse+Attention
12. FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference — Xunhao Lai, Jianqiao Lu, Yao Luo, Yiyuan Ma, Xun Zhou, 2025
https://scholar.google.com/scholar?q=FlexPrefill%3A+A+Context-Aware+Sparse+Attention+Mechanism+for+Efficient+Long-Sequence+Inference
13. MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool — Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 2024
https://scholar.google.com/scholar?q=MemServe%3A+Context+Caching+for+Disaggregated+LLM+Serving+with+Elastic+Memory+Pool
14. Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads — Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, Yizhou Shan, 2024
https://scholar.google.com/scholar?q=Inference+without+Interference%3A+Disaggregate+LLM+Inference+for+Mixed+Downstream+Workloads
15. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3
16. AI Post Transformers: Lookahead Q-Cache for Consistent KV Eviction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-lookahead-q-cache-for-consistent-kv-evic-d97b09.mp3
17. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
18. AI Post Transformers: ContiguousKV for Faster LLM Prefill KV Reuse — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-contiguouskv-for-faster-llm-prefill-kv-r-59f545.mp3
19. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
20. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
Interactive Visualization: ScoutAttention for Efficient KV Cache Offloading
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof