This episode explores Unified Neural Scaling Laws, a framework for predicting model performance when parameter count, data volume, training steps, inference compute, and training-recipe choices all change at once. It explains how the paper moves beyond classic smooth power-law curves by introducing broken multivariate scaling laws with regime shifts, including hyperbreaks, bottleneck versus non-bottleneck components, and joint interaction surfaces across training variables. The discussion highlights the paper’s argument that good forecasting must capture both beneficial scaling effects and harmful effects such as overfitting, bad hyperparameter regimes, and data or compute limits, rather than assuming one clean trend forever. A listener would find it interesting because it ties abstract scaling-law math directly to expensive real-world training decisions and to the question of whether pretraining gains actually transfer to downstream benchmarks.
Sources:
1. Unified Neural Scaling Laws — Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish, 2026
http://arxiv.org/abs/2605.26248
2. A Constructive Prediction of the Generalization Error Across Scales — Jonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir Shavit, 2020
https://scholar.google.com/scholar?q=A+Constructive+Prediction+of+the+Generalization+Error+Across+Scales
3. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, et al., 2020
https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models
4. Training Compute-Optimal Large Language Models — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models
5. Unified Neural Scaling Laws — Ethan Caballero, Priyank Jaini, David Krueger, Irina Rish, 2026
https://scholar.google.com/scholar?q=Unified+Neural+Scaling+Laws
6. Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour — Priya Goyal, Piotr Dollar, Ross Girshick, Kaiming He, et al., 2017
https://scholar.google.com/scholar?q=Accurate%2C+Large+Minibatch+SGD%3A+Training+ImageNet+in+1+Hour
7. An Empirical Model of Large-Batch Training — Sam McCandlish, Jared Kaplan, Dario Amodei, OpenAI Dota Team, 2018
https://scholar.google.com/scholar?q=An+Empirical+Model+of+Large-Batch+Training
8. Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer — Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, et al., 2022
https://scholar.google.com/scholar?q=Tensor+Programs+V%3A+Tuning+Large+Neural+Networks+via+Zero-Shot+Hyperparameter+Transfer
9. Resolving Discrepancies in Compute-Optimal Scaling of Language Models — Tomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt, Yair Carmon, 2024
https://scholar.google.com/scholar?q=Resolving+Discrepancies+in+Compute-Optimal+Scaling+of+Language+Models
10. Reconciling modern machine learning practice and the bias-variance trade-off — Mikhail Belkin, Daniel Hsu, Siyuan Ma, Soumik Mandal, 2019
https://scholar.google.com/scholar?q=Reconciling+modern+machine+learning+practice+and+the+bias-variance+trade-off
11. Deep Double Descent: Where Bigger Models and More Data Hurt — Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, Ilya Sutskever, 2020
https://scholar.google.com/scholar?q=Deep+Double+Descent%3A+Where+Bigger+Models+and+More+Data+Hurt
12. Broken Neural Scaling Laws — Ethan Caballero, Kshitij Gupta, Irina Rish, David Krueger, 2023
https://scholar.google.com/scholar?q=Broken+Neural+Scaling+Laws
13. Scaling Data-Constrained Language Models — Niklas Muennighoff, Alexander M. Rush, Boaz Barak, et al., 2023
https://scholar.google.com/scholar?q=Scaling+Data-Constrained+Language+Models
14. Observational Scaling Laws and the Predictability of Language Model Performance — Yangjun Ruan, Chris J. Maddison, Tatsunori Hashimoto, 2024
https://scholar.google.com/scholar?q=Observational+Scaling+Laws+and+the+Predictability+of+Language+Model+Performance
15. Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic — Sachin Goyal et al., 2024
https://scholar.google.com/scholar?q=Scaling+Laws+for+Data+Filtering+--+Data+Curation+cannot+be+Compute+Agnostic
16. Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies — Zhengyu Chen et al., 2025
https://scholar.google.com/scholar?q=Revisiting+Scaling+Laws+for+Language+Models%3A+The+Role+of+Data+Quality+and+Training+Strategies
17. Scaling Laws for Optimal Data Mixtures — Mustafa Shukor et al., 2025
https://scholar.google.com/scholar?q=Scaling+Laws+for+Optimal+Data+Mixtures
18. Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient — Jan Ludziejewski et al., 2025
https://scholar.google.com/scholar?q=Joint+MoE+Scaling+Laws%3A+Mixture+of+Experts+Can+Be+Memory+Efficient
19. Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning — Wenkai Yang et al., 2025
https://scholar.google.com/scholar?q=Towards+Thinking-Optimal+Scaling+of+Test-Time+Compute+for+LLM+Reasoning
20. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025
https://scholar.google.com/scholar?q=Scaling+up+Test-Time+Compute+with+Latent+Reasoning%3A+A+Recurrent+Depth+Approach
21. Towards Robust Scaling Laws for Optimizers — Alexandra Volkova et al., 2026
https://scholar.google.com/scholar?q=Towards+Robust+Scaling+Laws+for+Optimizers
22. Neural Scaling Laws Rooted in the Data Distribution — Ari Brill, 2024
https://scholar.google.com/scholar?q=Neural+Scaling+Laws+Rooted+in+the+Data+Distribution
23. The Quantization Model of Neural Scaling — Eric J. Michaud et al., 2023
https://scholar.google.com/scholar?q=The+Quantization+Model+of+Neural+Scaling
24. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
25. AI Post Transformers: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-14-tmas-scaling-test-time-compute-with-mult-3abe7a.mp3
26. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3
Interactive Visualization: Unified Neural Scaling Laws Across Regimes