This episode explores the RAGEN-2 paper’s claim that agentic reinforcement learning can produce reasoning traces that look active and diverse while losing real dependence on the input. It explains the paper’s central distinction between ordinary entropy, which measures diversity within a single prompt, and template collapse, where traces across many different prompts become generic variations of the same pattern. The discussion also covers the proposed mutual-information-style monitoring approach, which rescoring traces against other prompts to test whether reasoning remains identifiable to its source, and links the failure mode to weak reward signal, PPO-style regularization, and sparse long-horizon feedback. Listeners would find it interesting because it reframes a core question in reasoning RL: not whether an agent looks busy, but whether its reasoning is still actually about the problem in front of it.
Sources:
1. RAGEN-2: Reasoning Collapse in Agentic RL — Zihan Wang, Chi Gui, Xing Jin, Qineng Wang, Licheng Liu, Kangrui Wang, Shiqi Chen, Linjie Li, Zhengyuan Yang, Pingyue Zhang, Yiping Lu, Jiajun Wu, Li Fei-Fei, Lijuan Wang, Yejin Choi, Manling Li, 2026
http://arxiv.org/abs/2604.06268
2. MINE: Mutual Information Neural Estimation — Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, R. Devon Hjelm, 2018
https://scholar.google.com/scholar?q=MINE%3A+Mutual+Information+Neural+Estimation
3. Representation Learning with Contrastive Predictive Coding — Aaron van den Oord, Yazhe Li, Oriol Vinyals, 2018
https://scholar.google.com/scholar?q=Representation+Learning+with+Contrastive+Predictive+Coding
4. Learning deep representations by mutual information estimation and maximization — R. Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Phil Bachman, Adam Trischler, Yoshua Bengio, 2019
https://scholar.google.com/scholar?q=Learning+deep+representations+by+mutual+information+estimation+and+maximization
5. RAGEN-2: Reasoning Collapse in Agentic RL — Zihan Wang, Chi Gui, Xing Jin, Qineng Wang, Manling Li, et al., 2026
https://scholar.google.com/scholar?q=RAGEN-2%3A+Reasoning+Collapse+in+Agentic+RL
6. Proximal Policy Optimization Algorithms — John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov, 2017
https://scholar.google.com/scholar?q=Proximal+Policy+Optimization+Algorithms
7. Signal-to-Noise Ratio Analysis of Policy Gradient Algorithms — John W. Roberts, Russ Tedrake, 2008
https://scholar.google.com/scholar?q=Signal-to-Noise+Ratio+Analysis+of+Policy+Gradient+Algorithms
8. Understanding the Impact of Entropy on Policy Optimization — Zafarali Ahmed, Nicolas Le Roux, Mohammad Norouzi, Dale Schuurmans, 2019
https://scholar.google.com/scholar?q=Understanding+the+Impact+of+Entropy+on+Policy+Optimization
9. Prioritized Experience Replay — Tom Schaul, John Quan, Ioannis Antonoglou, David Silver, 2016
https://scholar.google.com/scholar?q=Prioritized+Experience+Replay
10. DAPO: An Open-Source LLM Reinforcement Learning System at Scale — Qiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, et al., 2025
https://scholar.google.com/scholar?q=DAPO%3A+An+Open-Source+LLM+Reinforcement+Learning+System+at+Scale
11. No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping — Thanh-Long V. Le, Myeongho Jeon, Kim Vu, Viet Lai, Eunho Yang, 2025
https://scholar.google.com/scholar?q=No+Prompt+Left+Behind%3A+Exploiting+Zero-Variance+Prompts+in+LLM+Reinforcement+Learning+via+Entropy-Guided+Advantage+Shaping
12. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning — Zihan Wang et al., 2025
https://scholar.google.com/scholar?q=RAGEN%3A+Understanding+Self-Evolution+in+LLM+Agents+via+Multi-Turn+Reinforcement+Learning
13. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models — Ganqu Cui et al., 2025
https://scholar.google.com/scholar?q=The+Entropy+Mechanism+of+Reinforcement+Learning+for+Reasoning+Language+Models
14. ASTER: Agentic Scaling with Tool-integrated Extended Reasoning — Xuqin Zhang, Quan He, Zhenrui Zheng, Zongzhang Zhang, Xu He, and Dong Li, 2026
https://scholar.google.com/scholar?q=ASTER%3A+Agentic+Scaling+with+Tool-integrated+Extended+Reasoning
15. Demystifying Reinforcement Learning in Agentic Reasoning — Zhaochen Yu, Ling Yang, Jiaru Zou, Shuicheng Yan, and Mengdi Wang, 2025
https://scholar.google.com/scholar?q=Demystifying+Reinforcement+Learning+in+Agentic+Reasoning
16. The Price of Format: Diversity Collapse in LLMs — Longfei Yun, Chenyang An, Zilong Wang, Letian Peng, and Jingbo Shang, 2025
https://scholar.google.com/scholar?q=The+Price+of+Format%3A+Diversity+Collapse+in+LLMs
17. Revisiting Entropy in Reinforcement Learning for Large Reasoning Models — Renren Jin et al., 2025
https://scholar.google.com/scholar?q=Revisiting+Entropy+in+Reinforcement+Learning+for+Large+Reasoning+Models
18. ERPO: Token-Level Entropy-Regulated Policy Optimization for Large Reasoning Models — Song Yu and Li Li, 2026
https://scholar.google.com/scholar?q=ERPO%3A+Token-Level+Entropy-Regulated+Policy+Optimization+for+Large+Reasoning+Models
19. Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning — Chen Qian et al., 2025
https://scholar.google.com/scholar?q=Demystifying+Reasoning+Dynamics+with+Mutual+Information%3A+Thinking+Tokens+are+Information+Peaks+in+LLM+Reasoning
20. MITS: Enhanced Tree Search Reasoning for LLMs via Pointwise Mutual Information — Jiaxi Li et al., 2025
https://scholar.google.com/scholar?q=MITS%3A+Enhanced+Tree+Search+Reasoning+for+LLMs+via+Pointwise+Mutual+Information
21. On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning — Yifan Zhang et al., 2025
https://scholar.google.com/scholar?q=On+the+Design+of+KL-Regularized+Policy+Gradient+Algorithms+for+LLM+Reasoning
22. Rethinking KL Regularization in RLHF: From Value Estimation to Gradient Optimization — Kezhao Liu et al., 2025
https://scholar.google.com/scholar?q=Rethinking+KL+Regularization+in+RLHF%3A+From+Value+Estimation+to+Gradient+Optimization
23. What Makes a Reward Model a Good Teacher? An Optimization Perspective — Noam Razin et al., 2025
https://scholar.google.com/scholar?q=What+Makes+a+Reward+Model+a+Good+Teacher%3F+An+Optimization+Perspective
24. Accelerating RLHF Training with Reward Variance Increase — Zonglin Yang et al., 2025
https://scholar.google.com/scholar?q=Accelerating+RLHF+Training+with+Reward+Variance+Increase
25. Efficient RLVR Training via Weighted Mutual Information Data Selection — Xinyu Zhou et al., 2026
https://scholar.google.com/scholar?q=Efficient+RLVR+Training+via+Weighted+Mutual+Information+Data+Selection
26. AI Post Transformers: Stabilizing Efficient Reasoning with Step-Level Advantage Selection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stabilizing-efficient-reasoning-with-ste-1e589d.mp3
27. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
28. AI Post Transformers: Benchmarking Test-Time Scaling for General LLM Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-benchmarking-test-time-scaling-for-gener-8f14f9.mp3
29. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3