This episode explores JANUS, a systems approach to serving mixture-of-experts transformers efficiently by separating attention layers from expert layers instead of deploying the whole model as a single monolithic unit. It explains why MoE models can still be expensive and latency-prone in practice: even if only a few experts activate per token, the system must still manage large expert memory footprints, skewed expert demand, and strict token-level latency targets such as time per output token. The discussion focuses on JANUS’s core ideas, including separate GPU pools for attention and expert computation, an adaptive two-phase communication scheme that reduces cross-node messaging overhead, and SLO-aware scaling that adjusts attention and expert capacity independently. Listeners would find it interesting because it turns MoE inference from a simple “sparse compute saves money” story into a deeper argument about distributed systems design, load balancing, and the real bottlenecks that determine whether advanced models feel fast in production.
Sources:
1. Janus: Disaggregating Attention and Experts for Scalable MoE Inference — Zhexiang Zhang, Ye Wang, Yumiao Zhao, Jiayu Xiao, Qianjing Yang, Xiangyu Wang, Jingzhe Jiang, Qizhen Weng, Ruichuan Chen, Shaohuai Shi, Adel N. Toosi, Yin Chen, Minchen Yu, 2025
http://arxiv.org/abs/2512.13525
2. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity — William Fedus, Barret Zoph, Noam Shazeer, 2021
https://scholar.google.com/scholar?q=Switch+Transformers%3A+Scaling+to+Trillion+Parameter+Models+with+Simple+and+Efficient+Sparsity
3. FastMoE: A Fast Mixture-of-Expert Training System — Jiaao He, Jiezhong Qiu, Aohan Zeng, Zhilin Yang, Jidong Zhai, Jie Tang, 2021
https://scholar.google.com/scholar?q=FastMoE%3A+A+Fast+Mixture-of-Expert+Training+System
4. DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, Hao Zhang, 2024
https://scholar.google.com/scholar?q=DistServe%3A+Disaggregating+Prefill+and+Decoding+for+Goodput-optimized+Large+Language+Model+Serving
5. MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism — Ruidong Zhu, Ziheng Jiang, Chao Jin, Peng Wu, Cesar A. Stuardo, Dongyang Wang and others, 2025
https://scholar.google.com/scholar?q=MegaScale-Infer%3A+Serving+Mixture-of-Experts+at+Scale+with+Disaggregated+Expert+Parallelism
6. eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference — Suraiya Tairin, Shohaib Mahmud, Haiying Shen, Anand Iyer, 2025
https://scholar.google.com/scholar?q=eMoE%3A+Task-aware+Memory+Efficient+Mixture-of-Experts-Based+%28MoE%29+Model+Inference
7. SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference — Luchang Li, Dongfang Li, Bozhao Gong, Yu Zhang, 2026
https://scholar.google.com/scholar?q=SLO-Aware+Compute+Resource+Allocation+for+Prefill-Decode+Disaggregated+LLM+Inference
8. MoEless: Efficient MoE LLM Serving via Serverless Computing — Hanfei Yu, Bei Ouyang, Shwai He, Ang Li, Hao Wang, 2026
https://scholar.google.com/scholar?q=MoEless%3A+Efficient+MoE+LLM+Serving+via+Serverless+Computing
9. xDeepServe: Model-as-a-Service on Huawei CloudMatrix384 — Ao Xiao et al., 2025
https://scholar.google.com/scholar?q=xDeepServe%3A+Model-as-a-Service+on+Huawei+CloudMatrix384
10. Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling — Yan Li, Zhenyu Zhang, Zhengang Wang, Pengfei Chen, Pengfei Zheng, 2025
https://scholar.google.com/scholar?q=Semantic+Parallelism%3A+Redefining+Efficient+MoE+Inference+via+Model-Data+Co-Scheduling
11. GRACE-MoE: Grouping and Replication with Locality-Aware Routing for Efficient Distributed MoE Inference — Yu Han, Lehan Pan, Jie Peng, Ziyang Tao, Hanqi Zhu, Wuyang Zhang, Yanyong Zhang, 2025
https://scholar.google.com/scholar?q=GRACE-MoE%3A+Grouping+and+Replication+with+Locality-Aware+Routing+for+Efficient+Distributed+MoE+Inference
12. BurstGPT: A Real-world Workload Dataset to Optimize LLM Serving Systems — Yuxin Wang, Yuhan Chen, Zeyu Li, Xueze Kang, Yuchu Fang, Yeju Zhou, Yang Zheng, Zhenheng Tang, Xin He, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, Xiaowen Chu, 2024
https://scholar.google.com/scholar?q=BurstGPT%3A+A+Real-world+Workload+Dataset+to+Optimize+LLM+Serving+Systems
13. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
14. Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference — Ranggi Hwang et al., 2023/2024
https://arxiv.org/abs/2308.12066
15. HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference — Peng Tang et al., 2024
https://arxiv.org/abs/2411.01433
16. DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference — Yujie Zhang, Shivam Aggarwal, Tulika Mitra, 2025
https://arxiv.org/abs/2501.10375
17. AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding — Zikun Li et al., 2025
https://arxiv.org/abs/2501.12162
18. SLOs-Serve: Optimized Serving of Multi-SLO LLMs — Siyuan Chen et al., 2025
https://arxiv.org/abs/2504.08784
19. Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement — Tian Wu et al., 2025
https://arxiv.org/abs/2508.12851
20. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
21. AI Post Transformers: Splitwise: Phase-Split LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-26-splitwise-phase-split-llm-inference-e8945b.mp3
22. AI Post Transformers: Computation-Bandwidth-Memory Trade-offs for AI Infrastructure — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-computation-bandwidth-memory-trade-offs-a83f2b.mp3
23. AI Post Transformers: Why LLM Serving Needs Mathematical Optimization — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-why-llm-serving-needs-mathematical-optim-647fc6.mp3
24. AI Post Transformers: LAPS for Length-Aware LLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-laps-for-length-aware-llm-serving-0c6149.mp3
25. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
26. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
Interactive Visualization: JANUS for Scalable MoE Inference