AI Post Transformers

EMO: Emergent Modularity for Mixture-of-Experts


Listen Later

Hal Turing and Dr. Ada Shannon take a deep dive into EMO: Pretraining Mixture of Experts for Emergent Modularity, a May 7, 2026 paper by Ryan Wang and co-authors from UC Berkeley and the Allen Institute for AI. The episode centers on a practical deployment question: if a workload is mostly code, math, or biomed, why must operators keep an entire giant model in memory instead of loading only the relevant slice? They frame EMO against the broader rise of sparse Mixture-of-Experts systems and explain why industry progress on active-parameter efficiency is not the same as delivering clean, domain-specific modules that can stand on their own at inference time.
The discussion carefully separates standard MoE behavior from the stronger notion of modularity that EMO is targeting. Hal and Ada walk through how sparse-gated MoE and Switch Transformer style routing already allow different tokens to activate different experts, but argue that this still leaves deployment looking monolithic because the router makes local token-level decisions rather than exposing stable task-level components. A biology prompt can still scatter across a messy set of experts, and the next sentence may hit a different set entirely. The hosts use that distinction to unpack the paper’s core concepts: emergent modularity from unlabeled data, semantic expert specialization around meaningful domains like code or math, composable architecture, and the memory-accuracy frontier that determines whether smaller loaded expert pools can preserve real capability.
The episode then gets into EMO’s training design and why the method is more than a single routing tweak. Ada explains the paper’s two-level routing scheme, where a document first selects a shared candidate pool of experts and individual tokens then choose active experts only within that pool, forcing document-consistent structure without removing all local flexibility. They also cover the supporting recipe: random pool-size sampling to expose the model to different memory budgets during training, global load balancing so a few experts do not dominate usage, and document-length-aware training so very long documents do not overwhelm the learning signal. The result is a focused discussion of whether MoE pretraining can produce expert groups that are not just sparsely activated, but genuinely deployable as modular tools.
Sources:
1. EMO: Emergent Modularity for Mixture-of-Experts
https://allenai.org/papers/emo
2. DEMix Layers: Disentangling Domains for Modular Language Modeling — Suchin Gururangan, Mike Lewis, Ari Holtzman, Noah A. Smith, Luke Zettlemoyer, 2021
https://scholar.google.com/scholar?q=DEMix+Layers%3A+Disentangling+Domains+for+Modular+Language+Modeling
3. ModuleFormer: Modularity Emerges from Mixture-of-Experts — Yikang Shen, Zheyu Zhang, Tianyou Cao, Shawn Tan, Zhenfang Chen, Chuang Gan, 2023
https://scholar.google.com/scholar?q=ModuleFormer%3A+Modularity+Emerges+from+Mixture-of-Experts
4. OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models — Fuzhao Xue, Zian Zheng, Yao Fu, Jinjie Ni, Zangwei Zheng, Wangchunshu Zhou, Yang You, 2024
https://scholar.google.com/scholar?q=OpenMoE%3A+An+Early+Effort+on+Open+Mixture-of-Experts+Language+Models
5. EMO: Pretraining Mixture of Experts for Emergent Modularity — Ryan Wang, Akshita Bhagia, Sewon Min, 2026
https://scholar.google.com/scholar?q=EMO%3A+Pretraining+Mixture+of+Experts+for+Emergent+Modularity
6. FlexOlmo: Open Language Models for Flexible Data Use — Weijia Shi et al., 2025
https://scholar.google.com/scholar?q=FlexOlmo%3A+Open+Language+Models+for+Flexible+Data+Use
7. Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM — Sainbayar Sukhbaatar et al., 2024
https://scholar.google.com/scholar?q=Branch-Train-MiX%3A+Mixing+Expert+LLMs+into+a+Mixture-of-Experts+LLM
8. The Myth of Expert Specialization in MoEs: Why Routing Reflects Geometry, Not Necessarily Domain Expertise — Xi Wang, Soufiane Hayou, Eric Nalisnick, 2026
https://scholar.google.com/scholar?q=The+Myth+of+Expert+Specialization+in+MoEs%3A+Why+Routing+Reflects+Geometry%2C+Not+Necessarily+Domain+Expertise
9. Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations — Zican Dong et al., 2025
https://scholar.google.com/scholar?q=Domain-Specific+Pruning+of+Large+Mixture-of-Experts+Models+with+Few-shot+Demonstrations
10. Efficient Expert Pruning for Sparse Mixture-of-Experts Language Models: Enhancing Performance and Reducing Inference Costs — Enshu Liu et al., 2024
https://scholar.google.com/scholar?q=Efficient+Expert+Pruning+for+Sparse+Mixture-of-Experts+Language+Models%3A+Enhancing+Performance+and+Reducing+Inference+Costs
11. MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router — Yanyue Xie et al., 2024
https://scholar.google.com/scholar?q=MoE-Pruner%3A+Pruning+Mixture-of-Experts+Large+Language+Model+using+the+Hints+from+Its+Router
12. AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models — Zihao Zeng et al., 2024
https://scholar.google.com/scholar?q=AdaMoE%3A+Token-Adaptive+Routing+with+Null+Experts+for+Mixture-of-Experts+Language+Models
13. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models — Damai Dai et al., 2024
https://scholar.google.com/scholar?q=DeepSeekMoE%3A+Towards+Ultimate+Expert+Specialization+in+Mixture-of-Experts+Language+Models
14. Fast Inference of Mixture-of-Experts Language Models with Offloading — Artyom Eliseev and Denis Mazur, 2023
https://scholar.google.com/scholar?q=Fast+Inference+of+Mixture-of-Experts+Language+Models+with+Offloading
15. Accelerating Mixture-of-Experts Inference by Hiding Offloading Latency with Speculative Decoding — Zhibin Wang et al., 2025
https://scholar.google.com/scholar?q=Accelerating+Mixture-of-Experts+Inference+by+Hiding+Offloading+Latency+with+Speculative+Decoding
16. AI Post Transformers: EMO: Emergent Modularity in Sparse Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-06-emo-emergent-modularity-in-sparse-langua-9551c4.mp3
17. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
18. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
19. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
Interactive Visualization: EMO: Emergent Modularity for Mixture-of-Experts
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof