This episode explores a 2025 survey of reinforcement learning as a statement about how the field now organizes itself, covering value-based, policy-based, model-based, multi-agent, offline, and LLM-related RL. It explains core concepts like Markov decision processes, policies, value functions, delayed credit assignment, and the contrast between direct policy optimization and methods that estimate action values before deriving behavior. The discussion highlights why actor-critic methods became so central, how model-based RL uses world models to plan ahead, and why offline RL is difficult when agents must improve from fixed logged data rather than fresh interaction. Listeners would find it interesting because it turns a broad survey into a clear map of where reinforcement learning stands in 2025, including the tensions between elegant theory, unstable training, and the practical compromises that shaped modern RL.
Sources:
1. Reinforcement Learning: An Overview — Kevin Murphy, 2024
http://arxiv.org/abs/2412.05265
2. Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems — Sergey Levine, Aviral Kumar, George Tucker, Justin Fu, 2020
https://scholar.google.com/scholar?q=Offline+Reinforcement+Learning%3A+Tutorial%2C+Review%2C+and+Perspectives+on+Open+Problems
3. D4RL: Datasets for Deep Data-Driven Reinforcement Learning — Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, Sergey Levine, 2020
https://scholar.google.com/scholar?q=D4RL%3A+Datasets+for+Deep+Data-Driven+Reinforcement+Learning
4. Conservative Q-Learning for Offline Reinforcement Learning — Aviral Kumar, Aurick Zhou, George Tucker, Sergey Levine, 2020
https://scholar.google.com/scholar?q=Conservative+Q-Learning+for+Offline+Reinforcement+Learning
5. Decision Transformer: Reinforcement Learning via Sequence Modeling — Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Aravind Srinivas, Igor Mordatch, 2021
https://scholar.google.com/scholar?q=Decision+Transformer%3A+Reinforcement+Learning+via+Sequence+Modeling
6. Reinforcement Learning: An Introduction — Richard S. Sutton and Andrew G. Barto, 2018
https://scholar.google.com/scholar?q=Reinforcement+Learning%3A+An+Introduction
7. Algorithms for Reinforcement Learning — Csaba Szepesvari, 2010
https://scholar.google.com/scholar?q=Algorithms+for+Reinforcement+Learning
8. Human-level control through deep reinforcement learning — Volodymyr Mnih, Koray Kavukcuoglu, David Silver and others, 2015
https://scholar.google.com/scholar?q=Human-level+control+through+deep+reinforcement+learning
9. Trust Region Policy Optimization — John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan and Pieter Abbeel, 2015
https://scholar.google.com/scholar?q=Trust+Region+Policy+Optimization
10. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
11. Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model — Julian Schrittwieser, Ioannis Antonoglou, Thomas Hubert and others, 2020
https://scholar.google.com/scholar?q=Mastering+Atari%2C+Go%2C+Chess+and+Shogi+by+Planning+with+a+Learned+Model
12. Fine-Tuning Language Models from Human Preferences — Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu and others, 2019
https://scholar.google.com/scholar?q=Fine-Tuning+Language+Models+from+Human+Preferences
13. Direct Preference Optimization: Your Language Model is Secretly a Reward Model — Rafael Rafailov, Archit Sharma, Eric Mitchell and others, 2023
https://scholar.google.com/scholar?q=Direct+Preference+Optimization%3A+Your+Language+Model+is+Secretly+a+Reward+Model
14. On the Limited Generalization Capability of the Implicit Reward Model Induced by Direct Preference Optimization — Yong Lin et al., 2024
https://scholar.google.com/scholar?q=On+the+Limited+Generalization+Capability+of+the+Implicit+Reward+Model+Induced+by+Direct+Preference+Optimization
15. Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RL — Taku Yamagata, Ahmed Khalil, Raul Santos-Rodriguez, 2023
https://scholar.google.com/scholar?q=Q-learning+Decision+Transformer%3A+Leveraging+Dynamic+Programming+for+Conditional+Sequence+Modelling+in+Offline+RL
16. Reinformer: Max-Return Sequence Modeling for Offline RL — Zifeng Zhuang et al., 2024
https://scholar.google.com/scholar?q=Reinformer%3A+Max-Return+Sequence+Modeling+for+Offline+RL
17. Pre-training Contextualized World Models with In-the-wild Videos for Reinforcement Learning — Jialong Wu, Haoyu Ma, Chaoyi Deng, Mingsheng Long, 2023
https://scholar.google.com/scholar?q=Pre-training+Contextualized+World+Models+with+In-the-wild+Videos+for+Reinforcement+Learning
18. PreLAR: World Model Pre-training with Learnable Action Representation — Lixuan Zhang, Meina Kan, Shiguang Shan, Xilin Chen, 2024
https://scholar.google.com/scholar?q=PreLAR%3A+World+Model+Pre-training+with+Learnable+Action+Representation
19. Ctrl-World: A Controllable Generative World Model for Robot Manipulation — Yanjiang Guo, Lucy Xiaoyang Shi, Jianyu Chen, Chelsea Finn, 2025
https://scholar.google.com/scholar?q=Ctrl-World%3A+A+Controllable+Generative+World+Model+for+Robot+Manipulation
20. Skill Transfer and Discovery for Sim-to-Real Learning: A Representation-Based Viewpoint — Haitong Ma, Zhaolin Ren, Bo Dai, Na Li, 2024
https://scholar.google.com/scholar?q=Skill+Transfer+and+Discovery+for+Sim-to-Real+Learning%3A+A+Representation-Based+Viewpoint
21. AI Post Transformers: DreamerV3 World Models Across 150 Tasks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-dreamerv3-world-models-across-150-tasks-af5edb.mp3
22. AI Post Transformers: Experience-Based Learning Beyond Human Data — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-experience-based-learning-beyond-human-d-b0caa4.mp3
23. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
24. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
Interactive Visualization: Reinforcement Learning in 2025: An Overview