AI Post Transformers

Moebius: Seamless Parallelism Switching for MoE Serving


Listen Later

This episode explores Moebius, a serving system for mixture-of-experts transformers that can switch at runtime between tensor parallelism and expert parallelism without restarting or draining live requests. It explains why tensor parallelism tends to give lower latency at low concurrency, while expert parallelism delivers better throughput at high concurrency, making bursty online traffic and RL rollouts natural settings where the best strategy changes over time. The discussion focuses on the hard systems problems behind that switch, including migrating in-flight requests, preserving paged KV caches, coping with CUDA graph address constraints, and handling KV-head mismatches that can waste cache capacity under tensor parallelism. It argues that the paper’s key contribution is treating the switch as a change in ownership and memory layout over one resident model and KV state, offering a concrete blueprint for serving large sparse models more efficiently.
Sources:
1. Moebius: Serving Mixture-of-Expert Models with Seamless Runtime Parallelism Switch — Shaoyu Wang, Yizhuo Liang, Jaeyong Song, Chong Li, Seo Jin Park, 2026
http://arxiv.org/abs/2606.26607
2. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding — Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, et al., 2020
https://scholar.google.com/scholar?q=GShard%3A+Scaling+Giant+Models+with+Conditional+Computation+and+Automatic+Sharding
3. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — William Fedus, Barret Zoph, Noam Shazeer, 2021
https://scholar.google.com/scholar?q=Switch+Transformers%3A+Scaling+to+Trillion+Parameter+Models+with+Simple+and+Efficient+Sparsity
4. DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale — Samyam Rajbhandari, Conglong Li, Zhewei Yao, et al., 2022
https://scholar.google.com/scholar?q=DeepSpeed-MoE%3A+Advancing+Mixture-of-Experts+Inference+and+Training+to+Power+Next-Generation+AI+Scale
5. MegaBlocks: Efficient Sparse Training with Mixture-of-Experts — Trevor Gale, Deepak Narayanan, Cliff Young, Matei Zaharia, 2022
https://scholar.google.com/scholar?q=MegaBlocks%3A+Efficient+Sparse+Training+with+Mixture-of-Experts
6. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism — Mohammad Shoeybi, Mostofa Patwary, Raul Puri, et al., 2019
https://scholar.google.com/scholar?q=Megatron-LM%3A+Training+Multi-Billion+Parameter+Language+Models+Using+Model+Parallelism
7. Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM — Deepak Narayanan, Mohammad Shoeybi, Jared Casper, et al., 2021
https://scholar.google.com/scholar?q=Efficient+Large-Scale+Language+Model+Training+on+GPU+Clusters+Using+Megatron-LM
8. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
9. Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism — Vikranth Srivatsa, Zijian He, Pu Guo, et al., 2026
https://scholar.google.com/scholar?q=Nitsum%3A+Serving+Tiered+LLM+Requests+with+Adaptive+Tensor+Parallelism
10. HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference — Haoran Lin et al., 2025
https://scholar.google.com/scholar?q=HAP%3A+Hybrid+Adaptive+Parallelism+for+Efficient+Mixture-of-Experts+Inference
11. Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services — Haoyu Chen et al., 2026
https://scholar.google.com/scholar?q=Amoeba%3A+Runtime+Tensor+Parallel+Transformation+for+LLM+Inference+Services
12. UCCL-EP: Portable Expert-Parallel Communication — Ziming Mao et al., 2026
https://scholar.google.com/scholar?q=UCCL-EP%3A+Portable+Expert-Parallel+Communication
13. RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training — Wei Gao et al., 2026
https://scholar.google.com/scholar?q=RollPacker%3A+Mitigating+Long-Tail+Rollouts+for+Fast%2C+Synchronous+RL+Post-Training
14. HD-MoE: Hybrid and Dynamic Parallelism for Mixture-of-Expert LLMs with 3D Near-Memory Processing — Haochen Huang et al., 2025
https://arxiv.org/abs/2509.09420
15. fMoE: Fine-Grained Expert Offloading for Large Mixture-of-Experts Serving — Hanfei Yu et al., 2025
https://arxiv.org/abs/2502.05370
16. HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference — Peng Tang et al., 2024
https://arxiv.org/abs/2411.01433
17. CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving — Yuhan Liu et al., 2023
https://arxiv.org/abs/2310.07240
18. KVzip: Query-Agnostic KV Cache Compression with Context Reconstruction — Jang-Hyun Kim et al., 2025
https://arxiv.org/abs/2505.23416
19. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
20. AI Post Transformers: Serving MoE Models with Disaggregated Expert Parallelism — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-serving-moe-models-with-disaggregated-ex-6979d2.mp3
21. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
22. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
23. AI Post Transformers: Mooncake for KV Cache-Centric LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-mooncake-for-kv-cache-centric-llm-servin-1086d0.mp3
24. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof