This episode explores a 2026 paper on making reasoning models faster and cheaper by shortening their chain-of-thought without sacrificing too much accuracy. It explains the paper’s core claim that short-context RL post-training can itself push models toward concise reasoning, while also unpacking the instability this creates and how the proposed Step-Level Advantage Selection method is meant to stabilize training. The discussion places that idea in the broader arc from chain-of-thought prompting and self-consistency to today’s industry-facing reasoning-budget controls, framing efficient reasoning as a test-time compute management problem rather than a new model architecture. Listeners would find it interesting for its skeptical, engineering-focused look at whether shorter reasoning traces are a real advance or just a fragile optimization hidden behind benchmark gains.
Sources:
1. Stabilizing Efficient Reasoning with Step-Level Advantage Selection — Han Wang, Xiaodong Yu, Jialian Wu, Jiang Liu, Ximeng Sun, Mohit Bansal, Zicheng Liu, 2026
http://arxiv.org/abs/2604.24003
2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models
3. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
4. Let’s Sample Step by Step: Adaptive-Consistency for Efficient Reasoning and Coding with LLMs — Pranjal Aggarwal, Aman Madaan, Yiming Yang, Mausam, 2023
https://scholar.google.com/scholar?q=Let%E2%80%99s+Sample+Step+by+Step%3A+Adaptive-Consistency+for+Efficient+Reasoning+and+Coding+with+LLMs
5. Training Language Models to Reason Efficiently — Daman Arora, Andrea Zanette, 2025
https://scholar.google.com/scholar?q=Training+Language+Models+to+Reason+Efficiently
6. DeepScaleR: Effective RL Scaling of Reasoning Models via Iterative Context Lengthening — Michael Luo, Sijun Tan, Justin Wong, Xiaoxiang Shi, William Y. Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, Ion Stoica, 2025
https://scholar.google.com/scholar?q=DeepScaleR%3A+Effective+RL+Scaling+of+Reasoning+Models+via+Iterative+Context+Lengthening
7. L1: Controlling How Long a Reasoning Model Thinks with Reinforcement Learning — Pranjal Aggarwal, Sean Welleck, 2025
https://scholar.google.com/scholar?q=L1%3A+Controlling+How+Long+a+Reasoning+Model+Thinks+with+Reinforcement+Learning
8. ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning — Bairu Hou, Yang Zhang, Jiabao Ji, Yujian Liu, Kaizhi Qian, Jacob Andreas, Shiyu Chang, 2025
https://scholar.google.com/scholar?q=ThinkPrune%3A+Pruning+Long+Chain-of-Thought+of+LLMs+via+Reinforcement+Learning
9. LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization — Xingyu Wu, Yuchen Yan, Shangke Lyu, Linjuan Wu, Yiwen Qiu, Yongliang Shen, Weiming Lu, Jian Shao, Jun Xiao, Yueting Zhuang, 2025
https://scholar.google.com/scholar?q=LAPO%3A+Internalizing+Reasoning+Efficiency+via+Length-Adaptive+Policy+Optimization
10. Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning — Shenzhi Wang, Le Yu, Chang Gao, Chujie Zheng, Shixuan Liu, Rui Lu, Kai Dang, Xiong-Hui Chen, Jianxin Yang, Zhenru Zhang, Yuqiong Liu, An Yang, Andrew Zhao, Yang Yue, Shiji Song, Bowen Yu, Gao Huang, Junyang Lin, 2025
https://scholar.google.com/scholar?q=Beyond+the+80%2F20+Rule%3A+High-Entropy+Minority+Tokens+Drive+Effective+Reinforcement+Learning+for+LLM+Reasoning
11. Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models — Xingyu Chen, Jiahao Xu, Tian Liang, Zhiwei He, Jianhui Pang, Dian Yu, Linfeng Song, Qiuzhi Liu, Mengfei Zhou, Zhuosheng Zhang, Rui Wang, Zhaopeng Tu, Haitao Mi, Dong Yu, 2025
https://scholar.google.com/scholar?q=Do+NOT+Think+That+Much+for+2%2B3%3D%3F+On+the+Overthinking+of+Long+Reasoning+Models
12. QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning — Fanqi Wan et al., 2025
https://scholar.google.com/scholar?q=QwenLong-L1%3A+Towards+Long-Context+Large+Reasoning+Models+with+Reinforcement+Learning
13. LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts — Siyuan Wang et al., 2025
https://scholar.google.com/scholar?q=LoongRL%3A+Reinforcement+Learning+for+Advanced+Reasoning+over+Long+Contexts
14. LongR: Unleashing Long-Context Reasoning via Reinforcement Learning with Dense Utility Rewards — Bowen Ping et al., 2026
https://scholar.google.com/scholar?q=LongR%3A+Unleashing+Long-Context+Reasoning+via+Reinforcement+Learning+with+Dense+Utility+Rewards
15. SSVPO: Effective Step-Level Credit Assignment for RL Training of Language Models — Yugu Li, Zehong Cao, Jianglin Qiao, Siyi Hu, 2026
https://scholar.google.com/scholar?q=SSVPO%3A+Effective+Step-Level+Credit+Assignment+for+RL+Training+of+Language+Models
16. GroundedPRM: Tree-Guided and Fidelity-Aware Process Reward Modeling for Step-Level Reasoning — Yao Zhang et al., 2025
https://scholar.google.com/scholar?q=GroundedPRM%3A+Tree-Guided+and+Fidelity-Aware+Process+Reward+Modeling+for+Step-Level+Reasoning
17. Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning — Miles Turpin, Andy Arditi, Marvin Li, Joe Benton, Julian Michael, 2025
https://scholar.google.com/scholar?q=Teaching+Models+to+Verbalize+Reward+Hacking+in+Chain-of-Thought+Reasoning
18. Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort — Xinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell, He He, 2025
https://scholar.google.com/scholar?q=Is+It+Thinking+or+Cheating%3F+Detecting+Implicit+Reward+Hacking+by+Measuring+Reasoning+Effort
19. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
20. AI Post Transformers: AgenticQwen and Small Industrial Tool Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-27-agenticqwen-and-small-industrial-tool-ag-dc676d.mp3
21. AI Post Transformers: World-R1 Improves 3D Consistency in Text-to-Video — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-28-world-r1-improves-3d-consistency-in-text-f065d9.mp3
22. AI Post Transformers: Experience-Based Learning Beyond Human Data — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-experience-based-learning-beyond-human-d-b0caa4.mp3
Interactive Visualization: Stabilizing Efficient Reasoning with Step-Level Advantage Selection