
Sign up to save your podcasts
Or


This episode is a deep dive into a post on Anthropic's research blog: https://www.anthropic.com/research/off-switch-dual-use
In this episode, James is joined by Ethan Roland, lead author of AE Studio's Gradient Routing paper, to explore a new approach to access control in frontier AI systems: modularizing dangerous capabilities during pre-training so they can be turned on and off at inference. The paper, "Modular Pre-Training Enables Access Control," was developed in collaboration with Anthropic, with roughly half of the co-authors coming from the lab.
Ethan makes the case for why current access control methods fall short. Inference-time guardrails get jailbroken in under 48 hours. Post hoc unlearning techniques like gradient ascent, RMU, and MaxEnt suppress capabilities superficially but let them snap back with 20 steps of fine-tuning. Data filtering is the gold standard, but naively requires training N separate frontier models to support N different dual-use categories, which is economically unworkable at hundreds of millions of dollars per pre-training run.
James and Ethan walk through how Gradient Routing solves this. A GR-MoE architecture uses one always-active core expert paired with smaller auxiliary experts, each responsible for a specific capability. During training, gradient updates from auxiliary-labeled data are frozen from touching the core, enforcing modularity at the parameter level. At inference, a binary configuration vector externally controls which auxiliaries participate. The method approximates the performance of full data filtering on both retained and ablated capabilities, but at the cost of a single pre-training run.
They also cover empirical results across scales from 50M to 2B parameters, the absorption effect that lets modularity persist under low labeling percentages, an arbitrary-subset variant that becomes exponentially more compute-efficient than data filtering, and how this connects to Andrej Karpathy's proposal for a cognitive core architecture in future AI systems.
Learn more: https://ae.studio/alignment
AE Studio is hiring: https://www.ae.studio/join-us
Subscribe to our newsletter: https://aestudio.beehiiv.com/
James Bowler LinkedIn: https://www.linkedin.com/in/james-bowler-84b02a100/
Ethan Roland LinkedIn: https://www.linkedin.com/in/ethan-roland/
Contact us: [email protected]
By James BowlerThis episode is a deep dive into a post on Anthropic's research blog: https://www.anthropic.com/research/off-switch-dual-use
In this episode, James is joined by Ethan Roland, lead author of AE Studio's Gradient Routing paper, to explore a new approach to access control in frontier AI systems: modularizing dangerous capabilities during pre-training so they can be turned on and off at inference. The paper, "Modular Pre-Training Enables Access Control," was developed in collaboration with Anthropic, with roughly half of the co-authors coming from the lab.
Ethan makes the case for why current access control methods fall short. Inference-time guardrails get jailbroken in under 48 hours. Post hoc unlearning techniques like gradient ascent, RMU, and MaxEnt suppress capabilities superficially but let them snap back with 20 steps of fine-tuning. Data filtering is the gold standard, but naively requires training N separate frontier models to support N different dual-use categories, which is economically unworkable at hundreds of millions of dollars per pre-training run.
James and Ethan walk through how Gradient Routing solves this. A GR-MoE architecture uses one always-active core expert paired with smaller auxiliary experts, each responsible for a specific capability. During training, gradient updates from auxiliary-labeled data are frozen from touching the core, enforcing modularity at the parameter level. At inference, a binary configuration vector externally controls which auxiliaries participate. The method approximates the performance of full data filtering on both retained and ablated capabilities, but at the cost of a single pre-training run.
They also cover empirical results across scales from 50M to 2B parameters, the absorption effect that lets modularity persist under low labeling percentages, an arbitrary-subset variant that becomes exponentially more compute-efficient than data filtering, and how this connects to Andrej Karpathy's proposal for a cognitive core architecture in future AI systems.
Learn more: https://ae.studio/alignment
AE Studio is hiring: https://www.ae.studio/join-us
Subscribe to our newsletter: https://aestudio.beehiiv.com/
James Bowler LinkedIn: https://www.linkedin.com/in/james-bowler-84b02a100/
Ethan Roland LinkedIn: https://www.linkedin.com/in/ethan-roland/
Contact us: [email protected]