This episode explores EMO, a Mixture-of-Experts language model that tries to turn token-level sparsity into real modularity by letting coherent expert groups emerge from document structure during pretraining. It explains why standard MoEs are not automatically deployable as smaller task-specific slices, then walks through EMO’s main idea: each document routes tokens through a learned document-specific pool of experts rather than the full expert set, with global load balancing to keep training stable. The discussion highlights that EMO was trained at substantial scale and reportedly matches a conventional MoE as a full model, while retaining surprisingly strong performance when only a fraction of experts are kept in memory for a given task. A listener would find it interesting because it connects a concrete systems problem, serving large models under tight memory budgets, to a plausible path toward more reusable, domain-specialized LLM components.
Sources:
1. EMO: Emergent Modularity in Sparse Language Models
https://allenai.org/papers/emo
2. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer — Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, Jeff Dean, 2017
https://scholar.google.com/scholar?q=Outrageously+Large+Neural+Networks%3A+The+Sparsely-Gated+Mixture-of-Experts+Layer
3. A Review of Sparse Expert Models in Deep Learning — William Fedus, Jeff Dean, Barret Zoph, 2022
https://scholar.google.com/scholar?q=A+Review+of+Sparse+Expert+Models+in+Deep+Learning
4. ST-MoE: Designing Stable and Transferable Sparse Expert Models — Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, William Fedus, 2022
https://scholar.google.com/scholar?q=ST-MoE%3A+Designing+Stable+and+Transferable+Sparse+Expert+Models
5. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models — Damai Dai, Chengqi Deng, Chenggang Zhao, Huazuo Gao, Deli Chen, Jiashi Li, Chong Ruan, Zhifang Sui, Wenfeng Liang, 2024
https://scholar.google.com/scholar?q=DeepSeekMoE%3A+Towards+Ultimate+Expert+Specialization+in+Mixture-of-Experts+Language+Models
6. Neural Module Networks — Jacob Andreas, Marcus Rohrbach, Trevor Darrell, Dan Klein, 2015
https://scholar.google.com/scholar?q=Neural+Module+Networks
7. AdapterFusion: Non-Destructive Task Composition for Transfer Learning — Jonas Pfeiffer, Aishwarya Kamath, Andreas Rucklé, Kyunghyun Cho, Iryna Gurevych, 2021
https://scholar.google.com/scholar?q=AdapterFusion%3A+Non-Destructive+Task+Composition+for+Transfer+Learning
8. Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners — Shashank Gupta, Subhabrata Mukherjee, Krishan Subudhi, Eduardo Gonzalez, Damien Jose, Ahmed H. Awadallah, Jianfeng Gao, 2022
https://scholar.google.com/scholar?q=Sparsely+Activated+Mixture-of-Experts+are+Robust+Multi-Task+Learners
9. GLaM: Efficient Scaling of Language Models with Mixture-of-Experts — Nan Du, Yanping Huang, Andrew Dai, Sebastian Goodman, Orhan Firat, Quoc Le, Yonghui Wu, Zhifeng Chen, Claire Cui, et al., 2021
https://scholar.google.com/scholar?q=GLaM%3A+Efficient+Scaling+of+Language+Models+with+Mixture-of-Experts
10. OLMoE: Open Mixture-of-Experts Language Models — Niklas Muennighoff et al., 2024
https://scholar.google.com/scholar?q=OLMoE%3A+Open+Mixture-of-Experts+Language+Models
11. Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM — Sainbayar Sukhbaatar et al., 2024
https://scholar.google.com/scholar?q=Branch-Train-MiX%3A+Mixing+Expert+LLMs+into+a+Mixture-of-Experts+LLM
12. FlexOlmo: Open Language Models for Flexible Data Use — Weijia Shi et al., 2025
https://scholar.google.com/scholar?q=FlexOlmo%3A+Open+Language+Models+for+Flexible+Data+Use
13. Emergent Modularity in Pre-trained Transformers — Zhengyan Zhang et al., 2023
https://scholar.google.com/scholar?q=Emergent+Modularity+in+Pre-trained+Transformers
14. Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations — Zican Dong et al., 2025
https://scholar.google.com/scholar?q=Domain-Specific+Pruning+of+Large+Mixture-of-Experts+Models+with+Few-shot+Demonstrations
15. Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models — Xudong Lu et al., 2024
https://arxiv.org/abs/2402.14800
16. Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs — Enshu Liu et al., 2024
https://arxiv.org/abs/2407.00945
17. The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise — Xi Wang, Soufiane Hayou, Eric Nalisnick, 2026
https://arxiv.org/abs/2604.09780
18. Dynamic Expert Specialization: Towards Catastrophic Forgetting-Free Multi-Domain MoE Adaptation — Junzhuo Li et al., 2025
https://arxiv.org/abs/2509.16882
19. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
20. AI Post Transformers: JANUS for Scalable MoE Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-janus-for-scalable-moe-inference-78ae30.mp3
21. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3