TalkRL: The Reinforcement Learning Podcast

TalkRL: The Reinforcement Learning Podcast

By Robin Ranjit Singh ChauhanTechnology
Download on the App Store

TalkRL: The Reinforcement Learning Podcast episodes

  • Finale Doshi-Velez on RL for Healthcare @ RCL 2024

    Finale Doshi-Velez is a Professor at the Harvard Paulson School of Engineering and Applied Sciences. 

    This off-the-cuff interview was recorded at UMass Amherst during the workshop day of RL Conference on August 9th 2024.   

    Host notes: I've been a fan of some of Prof Doshi-Velez' past work on clinical RL and hoped to feature her for some time now, so I jumped at the chance to get a few minutes of her thoughts -- even though you can tell I was not prepared and a bit flustered tbh.  Thanks to Prof Doshi-Velez for taking a moment for this, and I hope to cross paths in future for a more in depth interview.

    References  

    • Finale Doshi-Velez Homepage @ Harvard  
    • Finale Doshi-Velez on Google Scholar  


    8 min
  • David Silver 2 - Discussion after Keynote @ RCL 2024

    Thanks to Professor Silver for permission to record this discussion after his RLC 2024 keynote lecture.   

    Recorded at UMass Amherst during RCL 2024.

    Due to the live recording environment, audio quality varies.  We publish this audio in its raw form to preserve the authenticity and immediacy of the discussion.   

    References  

    • AlphaProof announcement on DeepMind's blog
    • Discovering Reinforcement Learning Algorithms, Oh et al  -- His keynote at RLC 2024 referred to more recent update to this work, yet to be published  
    • Reinforcement Learning Conference 2024  
    • David Silver on Google Scholar  
    17 min
  • David Silver @ RCL 2024

    David Silver is a principal research scientist at DeepMind and a professor at University College London. 

    This interview was recorded at UMass Amherst during RLC 2024.   

    References  

    • Discovering Reinforcement Learning Algorithms, Oh et al  -- His keynote at RLC 2024 referred to more recent update to this work, yet to be published  
    • Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm, Silver et al 2017 -- the AlphaZero algo was used   in his recent work on AlphaProof  
    • AlphaProof on the DeepMind blog 
    • AlphaFold on the DeepMind blog 
    • Reinforcement Learning Conference 2024  
    • David Silver on Google Scholar  
    12 min
  • Vincent Moens on TorchRL

    Dr. Vincent Moens is an Applied Machine Learning Research Scientist at Meta, and an author of TorchRL and TensorDict in pytorch. 

    Featured References

    TorchRL: A data-driven decision-making library for PyTorch
    Albert Bou, Matteo Bettini, Sebastian Dittert, Vikash Kumar, Shagun Sodhani, Xiaomeng Yang, Gianni De Fabritiis, Vincent Moens 


    Additional References  

    • TorchRL on github  
    • TensorDict Documentation  


    41 min
  • Arash Ahmadian on Rethinking RLHF

    Arash Ahmadian is a Researcher at Cohere and Cohere For AI focussed on Preference Training of large language models. He’s also a researcher at the Vector Institute of AI.

    Featured Reference

    Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

    Arash Ahmadian, Chris Cremer, Matthias Gallé, Marzieh Fadaee, Julia Kreutzer, Olivier Pietquin, Ahmet Üstün, Sara Hooker


    Additional References

    • Self-Rewarding Language Models, Yuan et al 2024
    • Reinforcement Learning: An Introduction, Sutton and Barto 1992
    • Learning from Delayed Rewards, Chris Watkins 1989
    • Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning, Williams 1992
    34 min
  • Glen Berseth on RL Conference

    Glen Berseth is an assistant professor at the Université de Montréal, a core academic member of the Mila - Quebec AI Institute, a Canada CIFAR AI chair, member l'Institute Courtios, and co-director of the Robotics and Embodied AI Lab (REAL). 

    Featured Links 

    Reinforcement Learning Conference 

    Closing the Gap between TD Learning and Supervised Learning--A Generalisation Point of View
    Raj Ghugare, Matthieu Geist, Glen Berseth, Benjamin Eysenbach

    22 min
  • Ian Osband

    Ian Osband is a Research scientist at OpenAI (ex DeepMind, Stanford) working on decision making under uncertainty.  

    We spoke about: 

    - Information theory and RL 

    - Exploration, epistemic uncertainty and joint predictions 

    - Epistemic Neural Networks and scaling to LLMs 


    Featured References 

    Reinforcement Learning, Bit by Bit 
    Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, Zheng Wen 

    From Predictions to Decisions: The Importance of Joint Predictive Distributions 

    Zheng Wen, Ian Osband, Chao Qin, Xiuyuan Lu, Morteza Ibrahimi, Vikranth Dwaracherla, Mohammad Asghari, Benjamin Van Roy  

     

    Epistemic Neural Networks 

    Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, Benjamin Van Roy  


    Approximate Thompson Sampling via Epistemic Neural Networks 

    Ian Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla, Morteza Ibrahimi, Xiuyuan Lu, Benjamin Van Roy 

      


    Additional References  

    • Thesis defence, Ian Osband 
    • Homepage, Ian Osband 
    • Epistemic Neural Networks at Stanford RL Forum 
    • Behaviour Suite for Reinforcement Learning, Osband et al 2019 
    • Efficient Exploration for LLMs, Dwaracherla et al 2024 
    1 hr 9 min
  • Sharath Chandra Raparthy

    Sharath Chandra Raparthy on In-Context Learning for Sequential Decision Tasks, GFlowNets, and more!  

    Sharath Chandra Raparthy is an AI Resident at FAIR at Meta, and did his Master's at Mila.  


    Featured Reference 

    Generalization to New Sequential Decision Making Tasks with In-Context Learning   
    Sharath Chandra Raparthy , Eric Hambro, Robert Kirk , Mikael Henaff, , Roberta Raileanu 

    Additional References  

    • Sharath Chandra Raparthy Homepage  
    • Human-Timescale Adaptation in an Open-Ended Task Space, Adaptive Agent Team 2023
    • Data Distributional Properties Drive Emergent In-Context Learning in Transformers, Chan et al 2022  
    • Decision Transformer: Reinforcement Learning via Sequence Modeling, Chen et al  2021


    41 min
  • Pierluca D'Oro and Martin Klissarov

    Pierluca D'Oro and Martin Klissarov on Motif and RLAIF, Noisy Neighborhoods and Return Landscapes, and more!  

    Pierluca D'Oro is PhD student at Mila and visiting researcher at Meta.


    Martin Klissarov is a PhD student at Mila and McGill and research scientist intern at Meta.  


    Featured References 

    Motif: Intrinsic Motivation from Artificial Intelligence Feedback 
    Martin Klissarov*, Pierluca D'Oro*, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, Mikael Henaff 

    Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control 
    Nate Rahn*, Pierluca D'Oro*, Harley Wiltzer, Pierre-Luc Bacon, Marc G. Bellemare 

    To keep doing RL research, stop calling yourself an RL researcher
    Pierluca D'Oro 

    58 min
  • Martin Riedmiller

    Martin Riedmiller of Google DeepMind on controlling nuclear fusion plasma in a tokamak with RL, the original Deep Q-Network, Neural Fitted Q-Iteration, Collect and Infer, AGI for control systems, and tons more!  


    Martin Riedmiller is a research scientist and team lead at DeepMind.   


    Featured References   


    Magnetic control of tokamak plasmas through deep reinforcement learning 
    Jonas Degrave, Federico Felici, Jonas Buchli, Michael Neunert, Brendan Tracey, Francesco Carpanese, Timo Ewalds, Roland Hafner, Abbas Abdolmaleki, Diego de las Casas, Craig Donner, Leslie Fritz, Cristian Galperti, Andrea Huber, James Keeling, Maria Tsimpoukelli, Jackie Kay, Antoine Merle, Jean-Marc Moret, Seb Noury, Federico Pesamosca, David Pfau, Olivier Sauter, Cristian Sommariva, Stefano Coda, Basil Duval, Ambrogio Fasoli, Pushmeet Kohli, Koray Kavukcuoglu, Demis Hassabis & Martin Riedmiller


    Human-level control through deep reinforcement learning
    Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, Demis Hassabis 

    Neural fitted Q iteration–first experiences with a data efficient neural reinforcement learning method 
    Martin Riedmiller  

    1 hr 14 min

About TalkRL: The Reinforcement Learning Podcast

From the publisher's feed

TalkRL podcast is All Reinforcement Learning, All the Time.

More shows like TalkRL: The Reinforcement Learning Podcast

Planet Money by NPR

Planet Money

30,701 Listeners

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,249 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,451 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,087 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas by Sean Carroll

Sean Carroll's Mindscape: Science, Society, Philosophy, Culture, Arts, and Ideas

4,164 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

203 Listeners

Google DeepMind: The Podcast by Hannah Fry

Google DeepMind: The Podcast

204 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

10,186 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

98 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

565 Listeners

Hard Fork by The New York Times

Hard Fork

5,557 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

140 Listeners

Latent Space: The AI Engineer Podcast by Latent.Space

Latent Space: The AI Engineer Podcast

102 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

684 Listeners