AI Post Transformers

LAPS for Length-Aware LLM Serving


Listen Later

This episode explores LAPS, a serving system for large language models that treats long prompt prefills and short multi-turn re-prefills as fundamentally different workloads instead of batching them together. It explains why user-perceived latency, especially time to first token, suffers when tiny follow-up requests get stuck behind large compute-heavy context loads, and how LAPS models the boundary between compute-bound and memory-bound prefills to separate them more intelligently. The discussion covers LAPS’s dual-queue design, its temporal and spatial disaggregation strategies, and engineering choices like short-request waiting windows, length-aware smart batching, and CUDA Graph execution. Listeners would find it interesting because it connects low-level scheduling and KV-cache behavior to the everyday experience of whether chat systems feel fast and responsive.
Sources:
1. LAPS: A Length-Aware-Prefill LLM Serving System — Jianshu She, Zonghang Li, Hongchao Du, Shangyu Wu, Wenhao Zheng, Eric Xing, Zhengzhong Liu, Huaxiu Yao, Jason Xue, Qirong Ho, 2026
http://arxiv.org/abs/2601.11589
2. ORCA: A Distributed Serving System for Transformer-Based Generative Models — Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, Byung-Gon Chun, 2022
https://scholar.google.com/scholar?q=ORCA%3A+A+Distributed+Serving+System+for+Transformer-Based+Generative+Models
3. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, Ion Stoica, 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
4. SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills — Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Ramachandran Ramjee, 2023
https://scholar.google.com/scholar?q=SARATHI%3A+Efficient+LLM+Inference+by+Piggybacking+Decodes+with+Chunked+Prefills
5. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
6. BucketServe: Bucket-Based Dynamic Batching for Smart and Efficient LLM Inference Serving — Wanyi Zheng, Minxian Xu, Shengye Song, Kejiang Ye, 2025
https://scholar.google.com/scholar?q=BucketServe%3A+Bucket-Based+Dynamic+Batching+for+Smart+and+Efficient+LLM+Inference+Serving
7. DéjàVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving — Foteini Strati, Sara McAllister, Amar Phanishayee, Jakub Tarnawski, Ana Klimovic, 2024
https://scholar.google.com/scholar?q=D%C3%A9j%C3%A0Vu%3A+KV-cache+Streaming+for+Fast%2C+Fault-tolerant+Generative+LLM+Serving
8. Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving — approx. contemporary LLM systems authors, 2025
https://scholar.google.com/scholar?q=Prefill-Decode+Aggregation+or+Disaggregation%3F+Unifying+Both+for+Goodput-Optimized+LLM+Serving
9. FlowPrefill: Decoupling Preemption from Prefill Scheduling Granularity to Mitigate Head-of-Line Blocking in LLM Serving — approx. contemporary LLM serving authors, 2025
https://scholar.google.com/scholar?q=FlowPrefill%3A+Decoupling+Preemption+from+Prefill+Scheduling+Granularity+to+Mitigate+Head-of-Line+Blocking+in+LLM+Serving
10. BurstGPT: A Real-World Workload Dataset to Optimize LLM Serving Systems — approx. systems/data-center workload authors, 2025
https://scholar.google.com/scholar?q=BurstGPT%3A+A+Real-World+Workload+Dataset+to+Optimize+LLM+Serving+Systems
11. SageServe: Optimizing LLM Serving on Cloud Data Centers with Forecast Aware Auto-Scaling — approx. cloud systems authors, 2025
https://scholar.google.com/scholar?q=SageServe%3A+Optimizing+LLM+Serving+on+Cloud+Data+Centers+with+Forecast+Aware+Auto-Scaling
12. ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production — approx. production-serving measurement authors, 2025
https://scholar.google.com/scholar?q=ServeGen%3A+Workload+Characterization+and+Generation+of+Large+Language+Model+Serving+in+Production
13. Fairness in Serving Large Language Models — approx. theory/systems fairness authors, 2025
https://scholar.google.com/scholar?q=Fairness+in+Serving+Large+Language+Models
14. FairBatching: Fairness-Aware Batch Formation for LLM Inference — approx. LLM inference scheduling authors, 2025
https://scholar.google.com/scholar?q=FairBatching%3A+Fairness-Aware+Batch+Formation+for+LLM+Inference
15. Locality-Aware Fair Scheduling in LLM Serving — approx. LLM serving systems authors, 2025
https://scholar.google.com/scholar?q=Locality-Aware+Fair+Scheduling+in+LLM+Serving
16. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
17. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
18. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
19. AI Post Transformers: Breaking the Prefix Barrier with Shared KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-breaking-the-prefix-barrier-with-shared-a5e5a6.mp3
20. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
Interactive Visualization: LAPS for Length-Aware LLM Serving
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof