AI Post Transformers

Stable Deep RL via Gaussian Representations


Listen Later

This episode explores a 2026 paper on stabilizing deep reinforcement learning by pushing an agent’s hidden representations toward an isotropic Gaussian shape. It explains how nonstationarity in RL, from shifting data distributions, bootstrapped targets, and primacy bias, can make agents overfit early experience, lose plasticity, and accumulate dormant neurons. The discussion focuses on the paper’s core argument that a round, evenly used feature space makes linear readouts easier to keep tracking as targets drift, reducing collapse and improving adaptation, and it breaks down SIGReg as a lightweight way to enforce that geometry. Listeners would find it interesting because it links an abstract idea from representation geometry to a concrete engineering problem in making RL systems more stable and trainable.
Sources:
1. Stable Deep Reinforcement Learning via Isotropic Gaussian Representations — Ali Saheb Pasand, Johan Obando-Ceron, Aaron Courville, Pouya Bashivan, Pablo Samuel Castro, 2026
http://arxiv.org/abs/2602.19373
2. Whitening for Self-Supervised Representation Learning — Aleksandr Ermolov, Aliaksandr Siarohin, Enver Sangineto, Nicu Sebe, 2020
https://scholar.google.com/scholar?q=Whitening+for+Self-Supervised+Representation+Learning
3. VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning — Adrien Bardes, Jean Ponce, Yann LeCun, 2021
https://scholar.google.com/scholar?q=VICReg%3A+Variance-Invariance-Covariance+Regularization+for+Self-Supervised+Learning
4. The Dormant Neuron Phenomenon in Deep Reinforcement Learning — Ghada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku Evci, 2023
https://scholar.google.com/scholar?q=The+Dormant+Neuron+Phenomenon+in+Deep+Reinforcement+Learning
5. LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics — Randall Balestriero, Yann LeCun, 2025
https://scholar.google.com/scholar?q=LeJEPA%3A+Provable+and+Scalable+Self-Supervised+Learning+Without+the+Heuristics
6. The Primacy Bias in Deep Reinforcement Learning — Evgenii Nikishin, Max Schwarzer, Pierluca D'Oro, Pierre-Luc Bacon, Aaron Courville, 2022
https://scholar.google.com/scholar?q=The+Primacy+Bias+in+Deep+Reinforcement+Learning
7. No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO — Skander Moalla, Andrea Miele, Daniil Pyatko, Razvan Pascanu, Caglar Gulcehre, 2024
https://scholar.google.com/scholar?q=No+Representation%2C+No+Trust%3A+Connecting+Representation%2C+Collapse%2C+and+Trust+Issues+in+PPO
8. Stable Gradients for Stable Learning at Scale in Deep Reinforcement Learning — Roger Creus Castanyer, Johan Obando-Ceron, Lu Li, Pierre-Luc Bacon, Glen Berseth, Aaron Courville, Pablo Samuel Castro, 2025
https://scholar.google.com/scholar?q=Stable+Gradients+for+Stable+Learning+at+Scale+in+Deep+Reinforcement+Learning
9. Proto-Value Networks: Scaling Representation Learning with Auxiliary Tasks — Jesse Farebrother et al., 2023
https://scholar.google.com/scholar?q=Proto-Value+Networks%3A+Scaling+Representation+Learning+with+Auxiliary+Tasks
10. Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics — Raphael Bernas et al., 2026
https://scholar.google.com/scholar?q=Revisiting+Anisotropy+in+Language+Transformers%3A+The+Geometry+of+Learning+Dynamics
11. Emergence of Quantised Representations Isolated to Anisotropic Functions — George Bird, 2025
https://scholar.google.com/scholar?q=Emergence+of+Quantised+Representations+Isolated+to+Anisotropic+Functions
12. AI Post Transformers: When Spectral Gradient Updates Help Deep Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-when-spectral-gradient-updates-help-deep-9c8441.mp3
13. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof