AI Post Transformers

IMPRESS: A Multi-Tier KV Storage System for Faster LLM Prefill


Listen Later

This episode explores IMPRESS, a systems paper from Zhejiang University and Huawei Cloud researchers presented at USENIX FAST 2025, which tackles a specific bottleneck in large language model inference: the time delay before a model produces its first response token when cached context has spilled onto disk. The discussion traces how modern LLM applications—retrieval-augmented generation, multi-turn chatbots, and plugin frameworks—prepend large chunks of context that inflate prefill costs superlinearly, with one cited example showing a 2,600-token plugin prompt stretching time-to-first-token by nine times on OPT-30B. It covers how prior work like vLLM's PagedAttention and AttentionStore addressed prefix KV caching but hit a wall once caches outgrew GPU and CPU memory and moved to slower disk storage, where I/O latency can consume up to 98 percent of total delay. The conversation traces IMPRESS's key insight: repurposing an importance-scoring technique from H2O—originally used to decide what to evict during decoding—to instead decide what's worth loading from disk before prefill even starts, claiming up to 2.8x lower latency with comparable accuracy. Listeners interested in the practical engineering tradeoffs behind making large-context LLM applications faster and cheaper to run will find the discussion's grounding in measured attention patterns, rather than benchmark tweaking, particularly compelling.
Sources:
1. IMPRESS: A Multi-Tier KV Storage System for Faster LLM Prefill
https://www.usenix.org/system/files/fast25-chen-weijian-impress.pdf
2. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica (UC Berkeley), 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
3. H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models — Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, Beidi Chen, 2023
https://scholar.google.com/scholar?q=H2O%3A+Heavy-Hitter+Oracle+for+Efficient+Generative+Inference+of+Large+Language+Models
4. Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention (AttentionStore) — Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, and colleagues (industry/academic collaboration), 2024
https://scholar.google.com/scholar?q=Cost-Efficient+Large+Language+Model+Serving+for+Multi-turn+Conversations+with+CachedAttention+%28AttentionStore%29
5. Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving — Ruoyu Qin, Zheming Li, Weiran He, Mingxing Zhang, Yongwei Wu, Weimin Zheng, Xinran Xu (Moonshot AI, Tsinghua University), 2024
https://scholar.google.com/scholar?q=Mooncake%3A+A+KVCache-centric+Disaggregated+Architecture+for+LLM+Serving
6. AttentionStore: Cost-Effective Attention Reuse across Multi-Turn Conversations in Large Language Model Serving — Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, Pengfei Zuo, 2024
https://scholar.google.com/scholar?q=AttentionStore%3A+Cost-Effective+Attention+Reuse+across+Multi-Turn+Conversations+in+Large+Language+Model+Serving
7. Retrieval Head Mechanistically Explains Long-Context Factuality — Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, Yao Fu, 2024
https://scholar.google.com/scholar?q=Retrieval+Head+Mechanistically+Explains+Long-Context+Factuality
8. RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation — Chao Jin, Zili Zhang, Xuanlin Jiang, Fangyue Liu, Xin Liu, Xuanzhe Liu, Xin Jin, 2024
https://scholar.google.com/scholar?q=RAGCache%3A+Efficient+Knowledge+Caching+for+Retrieval-Augmented+Generation
Interactive Visualization: IMPRESS: A Multi-Tier KV Storage System for Faster LLM Prefill
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof