AI Post Transformers

Serving MoE Models with Disaggregated Expert Parallelism


Listen Later

This episode explores MegaScale-Infer, a systems paper on serving large mixture-of-experts language models by separating the attention path from the expert feed-forward path and scheduling them independently. It explains why MoE models can look efficient on paper yet still waste GPU capacity in practice, especially during decode, where KV-cache-heavy attention and uneven expert routing create very different bottlenecks. The discussion focuses on the paper’s core argument for disaggregated expert parallelism and a ping-pong microbatch pipeline designed to keep both attention and expert GPUs busy instead of leaving one side idle. Listeners would find it interesting for its clear look at the gap between model architecture and real-world serving performance, including a pointed debate over whether strong decode benchmarks actually translate into better end-to-end user latency.
Sources:
1. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism — Ruidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo, Dongyang Wang, Xinlei Zhang, Huaping Zhou, Haoran Wei, Yang Cheng, Jianzhe Xiao, Xinyi Zhang, Lingjun Liu, Haibin Lin, Li-Wen Chang, Jianxi Ye, Xiao Yu, Xuanzhe Liu, Xin Jin, Xin Liu, 2025
http://arxiv.org/abs/2504.02263
2. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism — Yanping Huang, Yonglong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Miaosen Wang, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Zhifeng Chen, 2019
https://scholar.google.com/scholar?q=GPipe%3A+Efficient+Training+of+Giant+Neural+Networks+using+Pipeline+Parallelism
3. PipeDream: Generalized Pipeline Parallelism for DNN Training — Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, Phil Gibbons, Matei Zaharia, 2019
https://scholar.google.com/scholar?q=PipeDream%3A+Generalized+Pipeline+Parallelism+for+DNN+Training
4. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM — Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Nitin Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Matei Zaharia, 2021
https://scholar.google.com/scholar?q=Efficient+Large-Scale+Language+Model+Training+on+GPU+Clusters+Using+Megatron-LM
5. Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache — Bin Lin, Chen Zhang, Tao Peng, Hanyu Zhao, Wencong Xiao, Minmin Sun, Anmin Liu, Zhipeng Zhang, Lanbo Li, Xiafei Qiu, Shen Li, Zhigang Ji, Tao Xie, Yong Li, Wei Lin, 2024
https://scholar.google.com/scholar?q=Infinite-LLM%3A+Efficient+LLM+Service+for+Long+Context+with+DistAttention+and+Distributed+KVCache
6. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
7. Splitwise: Efficient Generative LLM Inference using Phase Splitting — Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Inigo Goiri, Saeed Maleki, Ricardo Bianchini, 2023
https://scholar.google.com/scholar?q=Splitwise%3A+Efficient+Generative+LLM+Inference+using+Phase+Splitting
8. MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs — Shiyi Cao, Shu Liu, Tyler Griggs, Peter Schafhalter, Xiaoxuan Liu, Ying Sheng, Joseph E. Gonzalez, Matei Zaharia, Ion Stoica, 2024
https://scholar.google.com/scholar?q=MoE-Lightning%3A+High-Throughput+MoE+Inference+on+Memory-constrained+GPUs
9. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
10. Toward Efficient Inference for Mixture of Experts — Haiyang Huang, Newsha Ardalani, Anna Sun, Liu Ke, Hsien-Hsin S. Lee, Shruti Bhosale, Carole-Jean Wu, Benjamin Lee, 2024
https://scholar.google.com/scholar?q=Toward+Efficient+Inference+for+Mixture+of+Experts
11. AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference — Shuzhang Zhong, Ling Liang, Yuan Wang, Runsheng Wang, Ru Huang, Meng Li, 2024
https://scholar.google.com/scholar?q=AdapMoE%3A+Adaptive+Sensitivity-based+Expert+Gating+and+Management+for+Efficient+MoE+Inference
12. HybriMoE: Hybrid CPU-GPU Scheduling and Cache Management for Efficient MoE Inference — Shuzhang Zhong, Yanfan Sun, Ling Liang, Runsheng Wang, Ru Huang, Meng Li, 2025
https://scholar.google.com/scholar?q=HybriMoE%3A+Hybrid+CPU-GPU+Scheduling+and+Cache+Management+for+Efficient+MoE+Inference
13. Oracle-MoE: Locality-preserving Routing in the Oracle Space for Memory-constrained Large Language Model Inference — Jixian Zhou, Fang Dong, Ruijun Huang, Hengjie Cao, Mengyi Chen, Yifeng Yang, Anrui Chen, Mingzhi Dong, Yujiang Wang, Dongsheng Li, David A. Clifton, Qin Lv, Rui Zhu, Chun Zhang, Fan Yang, Tun Lu, Ning Gu, Li Shang, 2025
https://scholar.google.com/scholar?q=Oracle-MoE%3A+Locality-preserving+Routing+in+the+Oracle+Space+for+Memory-constrained+Large+Language+Model+Inference
14. Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism — Xinglin Pan, Shaohuai Shi, Wenxiang Lin, Yuxin Wang, Zhenheng Tang, Wei Wang, Xiaowen Chu, 2025
https://scholar.google.com/scholar?q=Efficient+MoE+Inference+with+Fine-Grained+Scheduling+of+Disaggregated+Expert+Parallelism
15. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
16. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
17. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
18. AI Post Transformers: NanoFlow and the Future of LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-nanoflow-and-the-future-of-llm-serving-7429c9.mp3
19. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
20. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
21. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
22. AI Post Transformers: Deep Kernel Fusion for Transformer Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-deep-kernel-fusion-for-transformer-decod-b1a703.mp3
Interactive Visualization: Serving MoE Models with Disaggregated Expert Parallelism
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof