AI Post Transformers

Agentic Discovery for Test-Time Scaling


Listen Later

This episode explores a paper on test-time scaling that asks whether an LLM agent can automatically discover better inference-time control policies than the hand-built heuristics researchers usually rely on. It explains the core search framework in concrete terms: a controller decides when to branch, continue, probe, prune, or stop, balancing reasoning depth, breadth, and compute budget rather than simply generating more tokens. The discussion highlights the paper’s main technical argument that offline replay over logged reasoning traces, combined with a compact controller parameterization and detailed execution feedback, makes policy discovery cheap enough to be practical. Listeners would find it interesting because it connects abstract ideas about reasoning agents to a very specific claim: smarter inference may come not from larger models, but from better learned strategies for spending compute.
Sources:
1. LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling — Tong Zheng, Haolin Liu, Chengsong Huang, Huiwen Bao, Sheng Zhang, Rui Liu, Runpeng Dai, Ruibo Chen, Chenxi Liu, Tianyi Xiong, Xidong Wu, Hongming Zhang, Heng Huang, 2026
http://arxiv.org/abs/2605.08083
2. Supervisory Control of a Class of Discrete Event Systems — Peter J. Ramadge, Walter M. Wonham, 1987
https://scholar.google.com/scholar?q=Supervisory+Control+of+a+Class+of+Discrete+Event+Systems
3. On the Synthesis of a Reactive Module — Amir Pnueli, Roni Rosner, 1989
https://scholar.google.com/scholar?q=On+the+Synthesis+of+a+Reactive+Module
4. Formal Methods for Control Synthesis: An Optimization Perspective — Calin Belta, Sadra Sadraddini, 2019
https://scholar.google.com/scholar?q=Formal+Methods+for+Control+Synthesis%3A+An+Optimization+Perspective
5. Formal Synthesis of Controllers for Safety-Critical Autonomous Systems: Developments and Challenges — Xiang Yin, Bingzhao Gao, Xiao Yu, 2024
https://scholar.google.com/scholar?q=Formal+Synthesis+of+Controllers+for+Safety-Critical+Autonomous+Systems%3A+Developments+and+Challenges
6. Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems — Sergey Levine, Aviral Kumar, George Tucker, Justin Fu, 2020
https://scholar.google.com/scholar?q=Offline+Reinforcement+Learning%3A+Tutorial%2C+Review%2C+and+Perspectives+on+Open+Problems
7. Off-Policy Deep Reinforcement Learning without Exploration — Scott Fujimoto, David Meger, Doina Precup, 2019
https://scholar.google.com/scholar?q=Off-Policy+Deep+Reinforcement+Learning+without+Exploration
8. Conservative Q-Learning for Offline Reinforcement Learning — Aviral Kumar, Aurick Zhou, George Tucker, Sergey Levine, 2020
https://scholar.google.com/scholar?q=Conservative+Q-Learning+for+Offline+Reinforcement+Learning
9. D4RL: Datasets for Deep Data-Driven Reinforcement Learning — Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, Sergey Levine, 2020
https://scholar.google.com/scholar?q=D4RL%3A+Datasets+for+Deep+Data-Driven+Reinforcement+Learning
10. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally+can+be+More+Effective+than+Scaling+Model+Parameters
11. Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing — Tong Zheng, Chengsong Huang, Runpeng Dai, Yun He, Rui Liu, Xin Ni, Huiwen Bao, Kaishen Wang, Hongtu Zhu, Jiaxin Huang, Furong Huang, Heng Huang, 2026
https://scholar.google.com/scholar?q=Parallel-Probe%3A+Towards+Efficient+Parallel+Thinking+via+2D+Probing
12. Scaling Test-time Compute for LLM Agents — King Zhu, Hanhao Li, Siwei Wu, Tianshun Xing, Dehua Ma, Xiangru Tang, Minghao Liu, Jian Yang, Jiaheng Liu, Yuchen Eleanor Jiang, Changwang Zhang, Chenghua Lin, Jun Wang, Ge Zhang, Wangchunshu Zhou, 2025
https://scholar.google.com/scholar?q=Scaling+Test-time+Compute+for+LLM+Agents
13. Timely Machine: Awareness of Time Makes Test-Time Scaling Agentic — Yichuan Ma, Linyang Li, Yongkang Chen, Peiji Li, Xiaozhe Li, Qipeng Guo, Dahua Lin, Kai Chen, 2026
https://scholar.google.com/scholar?q=Timely+Machine%3A+Awareness+of+Time+Makes+Test-Time+Scaling+Agentic
14. Predicting and improving test-time scaling laws via reward tail-guided search — Muheng Li, Jian Qian, Wenlong Mou, 2026
https://scholar.google.com/scholar?q=Predicting+and+improving+test-time+scaling+laws+via+reward+tail-guided+search
15. TTRL: Test-Time Reinforcement Learning — Yuxin Zuo, Kaiyan Zhang, Shang Qu, Li Sheng, et al., 2025
https://scholar.google.com/scholar?q=TTRL%3A+Test-Time+Reinforcement+Learning
16. CTRLS: Chain-of-Thought Reasoning via Latent State-Transition — Junda Wu, Yuxin Xiong, Xintong Li, Zhengmian Hu, Tong Yu, Rui Wang, Xiang Chen, Jingbo Shang, Julian McAuley, 2025
https://scholar.google.com/scholar?q=CTRLS%3A+Chain-of-Thought+Reasoning+via+Latent+State-Transition
17. Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification — Anqi Zhang, Yulin Chen, Jane Pan, Chen Zhao, Aurojit Panda, Jinyang Li, He He, 2025
https://scholar.google.com/scholar?q=Reasoning+Models+Know+When+They%27re+Right%3A+Probing+Hidden+States+for+Self-Verification
18. Calibrated Reasoning: An Explanatory Verifier for Dynamic and Efficient Problem-Solving — Anisha Garg, Engin Tekin, Yash More, David Bick, Nishit Neema, Ganesh Venkatesh, 2025
https://scholar.google.com/scholar?q=Calibrated+Reasoning%3A+An+Explanatory+Verifier+for+Dynamic+and+Efficient+Problem-Solving
19. Contextual Drag: How Errors in the Context Affect LLM Reasoning — Yun Cheng, Xingyu Zhu, Haoyu Zhao, Sanjeev Arora, 2026
https://scholar.google.com/scholar?q=Contextual+Drag%3A+How+Errors+in+the+Context+Affect+LLM+Reasoning
20. Reinforcement Learning Teachers of Test Time Scaling — Edoardo Cetin, Tianyu Zhao, Yujin Tang, 2025
https://scholar.google.com/scholar?q=Reinforcement+Learning+Teachers+of+Test+Time+Scaling
21. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
22. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
23. AI Post Transformers: Generalist Reward Modeling with Inference-Time Scaling — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/generalist-reward-modeling-with-inference-time-scaling/
24. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
25. AI Post Transformers: SGLang for Faster Structured LLM Programs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-06-sglang-for-faster-structured-llm-program-c59f1c.mp3
Interactive Visualization: Agentic Discovery for Test-Time Scaling
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof