This episode explores a 2025 DeepMind and University of Alberta preprint arguing that AI is reaching the limits of learning from human-generated data and that the next major advances will come from agents learning through interaction and feedback in environments. It explains the shift from static pretraining to grounded reinforcement learning, defining key ideas like long-term reward optimization, self-play, world models, and why this paradigm has powered systems such as AlphaGo Zero, AlphaZero, MuZero, and theorem-proving agents in verifier-rich domains like math, code, and games. The discussion also stresses the practical obstacles that have kept RL from dominating mainstream AI—expensive data collection, sparse rewards, instability, and safety concerns—and questions whether this “era of experience” will extend broadly or remain strongest in environments where success can be automatically checked. Listeners would find it interesting for its clear breakdown of a major proposed shift in AI research and its skeptical take on whether the evidence really supports such a sweeping roadmap.
Sources:
1. Experience-Based Learning Beyond Human Data
https://storage.googleapis.com/deepmind-media/Era-of-Experience%20/The%20Era%20of%20Experience%20Paper.pdf
2. https://www.lesswrong.com/posts/TCGgiJAinGgcMEByt/the-era-of-experience-has-an-unsolved-technical-alignment
https://www.lesswrong.com/posts/TCGgiJAinGgcMEByt/the-era-of-experience-has-an-unsolved-technical-alignment
3. Reinforcement Learning: An Introduction — Richard S. Sutton, Andrew G. Barto, 1998; 2nd edition 2018
https://scholar.google.com/scholar?q=Reinforcement+Learning%3A+An+Introduction
4. Human-level control through deep reinforcement learning — Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu and others, 2015
https://scholar.google.com/scholar?q=Human-level+control+through+deep+reinforcement+learning
5. Deep Reinforcement Learning: An Overview — Yuxi Li, 2017
https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning%3A+An+Overview
6. Mastering Diverse Domains through World Models — Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert, Karen Simonyan, Laurent Sifre, Simon Schmitt, Arthur Guez, Edward Lockhart, Demis Hassabis, Thore Graepel, Timothy Lillicrap, David Silver, 2020
https://scholar.google.com/scholar?q=Mastering+Diverse+Domains+through+World+Models
7. AlphaProof and AlphaGeometry 2 — DeepMind et al., 2024
https://scholar.google.com/scholar?q=AlphaProof+and+AlphaGeometry+2
8. Mastering the Game of Go without Human Knowledge — David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George van den Driessche, Thore Graepel, Demis Hassabis, 2017
https://scholar.google.com/scholar?q=Mastering+the+Game+of+Go+without+Human+Knowledge
9. Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm — David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, Demis Hassabis, 2018
https://scholar.google.com/scholar?q=Mastering+Chess+and+Shogi+by+Self-Play+with+a+General+Reinforcement+Learning+Algorithm
10. Reward is Enough — David Silver, Satinder Singh, Doina Precup, Richard S. Sutton, 2021
https://scholar.google.com/scholar?q=Reward+is+Enough
11. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models — DeepSeek-AI et al., 2024
https://scholar.google.com/scholar?q=DeepSeekMath%3A+Pushing+the+Limits+of+Mathematical+Reasoning+in+Open+Language+Models
12. Reinforcement Learning from Human Feedback: Learning Dynamic Choices via Human Preferences — Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, Dario Amodei, 2017
https://scholar.google.com/scholar?q=Reinforcement+Learning+from+Human+Feedback%3A+Learning+Dynamic+Choices+via+Human+Preferences
13. Agent Lightning: Train ANY AI Agents with Reinforcement Learning — Yunfan Luo, et al., 2025
https://scholar.google.com/scholar?q=Agent+Lightning%3A+Train+ANY+AI+Agents+with+Reinforcement+Learning
14. Beyond human data: Scaling self-training for problem-solving with language models — approx. recent LLM self-training authors, 2024-2025
https://scholar.google.com/scholar?q=Beyond+human+data%3A+Scaling+self-training+for+problem-solving+with+language+models
15. Personalizing reinforcement learning from human feedback with variational preference learning — approx. recent preference-learning authors, 2024-2025
https://scholar.google.com/scholar?q=Personalizing+reinforcement+learning+from+human+feedback+with+variational+preference+learning
16. Online iterative reinforcement learning from human feedback with general preference model — approx. recent RLHF authors, 2024-2025
https://scholar.google.com/scholar?q=Online+iterative+reinforcement+learning+from+human+feedback+with+general+preference+model
17. Efficient preference-based reinforcement learning using learned dynamics models — approx. recent model-based preference RL authors, 2023-2025
https://scholar.google.com/scholar?q=Efficient+preference-based+reinforcement+learning+using+learned+dynamics+models
18. Refining Large Language Models with Self-Generated Data Through Iterative Training — approx. recent self-generated data / iterative training authors, 2024-2025
https://scholar.google.com/scholar?q=Refining+Large+Language+Models+with+Self-Generated+Data+Through+Iterative+Training
19. Co-evolved Self-Critique: Enhancing Large Language Models with Self-Generated Data — approx. recent self-critique authors, 2024-2025
https://scholar.google.com/scholar?q=Co-evolved+Self-Critique%3A+Enhancing+Large+Language+Models+with+Self-Generated+Data
20. Agentic reward modeling: Integrating human preferences with verifiable correctness signals for reliable reward systems — approx. recent reward-modeling authors, 2024-2025
https://scholar.google.com/scholar?q=Agentic+reward+modeling%3A+Integrating+human+preferences+with+verifiable+correctness+signals+for+reliable+reward+systems
21. Crossing the reward bridge: Expanding RL with verifiable rewards across diverse domains — approx. recent RLVR authors, 2024-2025
https://scholar.google.com/scholar?q=Crossing+the+reward+bridge%3A+Expanding+RL+with+verifiable+rewards+across+diverse+domains
22. AI Post Transformers: Experiential Reinforcement Learning: Internalizing Reflection for Better Policy Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/experiential-reinforcement-learning-internalizing-reflection-for-better-policy-t/
23. AI Post Transformers: MEMSEARCHER: Reinforcement Learning for LLM Memory Management — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-memsearcher-reinforcement-learning-for-l-e9ad84.mp3
24. AI Post Transformers: MetaClaw: Just Talk and Continual Agent Adaptation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-31-metaclaw-meta-learning-agents-in-the-wil-ab324c.mp3
25. AI Post Transformers: Memory Intelligence Agents for Deep Research — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-memory-intelligence-agents-for-deep-rese-cd39e3.mp3
26. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
27. AI Post Transformers: Recursive Language Models for Arbitrarily Long Prompts — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-recursive-language-models-for-arbitraril-fbcd1c.mp3
28. AI Post Transformers: Simple Self-Distillation for Better Code Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-02-simple-self-distillation-for-better-code-cc88e0.mp3
29. AI Post Transformers: IMO-Bench for Robust Mathematical Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-imo-bench-for-robust-mathematical-reason-143489.mp3
30. AI Post Transformers: AI Agent Traps and Prompt Injection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-02-ai-agent-traps-and-prompt-injection-7ce4ba.mp3
31. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
Interactive Visualization: Experience-Based Learning Beyond Human Data