This episode explores TMAS, a framework for scaling test-time reasoning by coordinating multiple specialized agents instead of simply letting a single model think longer. It explains how the system combines proposal, verification, refinement, and shared hierarchical memory so that useful intermediate results and higher-level strategy guidance can be reused across parallel reasoning attempts. The discussion highlights the paper’s central argument that better orchestration, careful memory design, and reinforcement learning objectives for exploration and productive memory use can turn extra inference compute into genuine reasoning gains rather than redundant or noisy work. Listeners would find it interesting for its clear comparison to self-consistency, Tree of Thoughts, and newer coordinated-reasoning methods, along with its emphasis on compute-matched evaluation, reproducibility, and the practical challenge of making multi-agent “synergy” real instead of just expensive parallelism.
Sources:
1. TMAS: Scaling Test-Time Compute via Multi-Agent Synergy — George Wu, Nan Jing, Qing Yi, Chuan Hao, Ming Yang, Feng Chang, Yuan Wei, Jian Yang, Ran Tao, Bryan Dai, 2026
http://arxiv.org/abs/2605.10344
2. A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well? — Qiyuan Zhang, Fuyuan Lyu, Zexu Sun, Lei Wang, Weixu Zhang, Wenyue Hua, Haolun Wu, Zhihan Guo, Yufei Wang, Niklas Muennighoff, Irwin King, Xue Liu, Chen Ma, 2025
https://scholar.google.com/scholar?q=A+Survey+on+Test-Time+Scaling+in+Large+Language+Models%3A+What%2C+How%2C+Where%2C+and+How+Well%3F
3. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2023
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
4. Tree of Thoughts: Deliberate Problem Solving with Large Language Models — Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, Karthik Narasimhan, 2023
https://scholar.google.com/scholar?q=Tree+of+Thoughts%3A+Deliberate+Problem+Solving+with+Large+Language+Models
5. PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning — Jingcheng Hu, Yinmin Zhang, Shijie Shang, Xiaobo Yang, Yue Peng, Zhewei Huang, Hebin Zhou, Xin Wu, Jie Cheng, Fanqi Wan, Xiangwen Kong, Chengyuan Yao, Kaiwen Yan, Ailin Huang, Hongyu Zhou, Qi Han, Zheng Ge, Daxin Jiang, Xiangyu Zhang, Heung-Yeung Shum, 2026
https://scholar.google.com/scholar?q=PaCoRe%3A+Learning+to+Scale+Test-Time+Compute+with+Parallel+Coordinated+Reasoning
6. Do Not Waste Your Rollouts: Recycling Search Experience for Efficient Test-Time Scaling — Xinglin Wang, Jiayi Shi, Shaoxiong Feng, Peiwen Yuan, Yiwei Li, Yueqi Zhang, Chuyi Tan, Ji Zhang, Boyuan Pan, Yao Hu, Kan Li, 2026
https://scholar.google.com/scholar?q=Do+Not+Waste+Your+Rollouts%3A+Recycling+Search+Experience+for+Efficient+Test-Time+Scaling
7. Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers — Shalev Lifshitz, Sheila A. McIlraith, Yilun Du, 2025
https://scholar.google.com/scholar?q=Multi-Agent+Verification%3A+Scaling+Test-Time+Compute+with+Multiple+Verifiers
8. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://scholar.google.com/scholar?q=Scaling+LLM+Test-Time+Compute+Optimally+can+be+More+Effective+than+Scaling+Model+Parameters
9. Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation — Zhuolin Yang, Zihan Liu, Yang Chen, Wenliang Dai, Boxin Wang, Sheng-Chieh Lin, Chankyu Lee, Yangyi Chen, Dongfu Jiang, Jiafan He, Renjie Pi, Grace Lam, Nayeon Lee, Alexander Bukharin, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping, 2026
https://scholar.google.com/scholar?q=Nemotron-Cascade+2%3A+Post-Training+LLMs+with+Cascade+RL+and+Multi-Domain+On-Policy+Distillation
10. ArcMemo: Abstract Reasoning Composition with Lifelong LLM Memory — Matthew Ho et al., 2025
https://scholar.google.com/scholar?q=ArcMemo%3A+Abstract+Reasoning+Composition+with+Lifelong+LLM+Memory
11. Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors — Aniket Didolkar, Nicolas Ballas, Sanjeev Arora, Anirudh Goyal, 2025
https://scholar.google.com/scholar?q=Metacognitive+Reuse%3A+Turning+Recurring+LLM+Reasoning+Into+Concise+Behaviors
12. Reuse, Don't Recompute: Efficient Large Reasoning Model Inference via Memory Orchestration — Daivik Patel, Shrenik Patel, 2025
https://scholar.google.com/scholar?q=Reuse%2C+Don%27t+Recompute%3A+Efficient+Large+Reasoning+Model+Inference+via+Memory+Orchestration
13. Diversity of Thought Improves Reasoning Abilities of Large Language Models — Ranjita Naik et al., 2023
https://scholar.google.com/scholar?q=Diversity+of+Thought+Improves+Reasoning+Abilities+of+Large+Language+Models
14. Diversity-Enhanced Reasoning for Subjective Questions — Yumeng Wang et al., 2025
https://scholar.google.com/scholar?q=Diversity-Enhanced+Reasoning+for+Subjective+Questions
15. Scaling Large Language Model-based Multi-Agent Collaboration — Chen Qian et al., 2024
https://scholar.google.com/scholar?q=Scaling+Large+Language+Model-based+Multi-Agent+Collaboration
16. Multi-Agent Sampling: Scaling Inference Compute for Data Synthesis with Tree Search-Based Agentic Collaboration — Hai Ye, Mingbao Lin, Hwee Tou Ng, Shuicheng Yan, 2024
https://scholar.google.com/scholar?q=Multi-Agent+Sampling%3A+Scaling+Inference+Compute+for+Data+Synthesis+with+Tree+Search-Based+Agentic+Collaboration
17. Reasoning Relay: Evaluating Stability and Interchangeability of Large Language Models in Mathematical Reasoning — Leo Lu et al., 2025
https://scholar.google.com/scholar?q=Reasoning+Relay%3A+Evaluating+Stability+and+Interchangeability+of+Large+Language+Models+in+Mathematical+Reasoning
18. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
19. AI Post Transformers: TUMIX Multi-Agent Test-Time Scaling with Tools — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-tumix-multi-agent-test-time-scaling-with-40671c.mp3
20. AI Post Transformers: DeepVerifier: Self-Evolving Research Agents via Rubric-Guided Verification — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/deepverifier-self-evolving-research-agents-via-rubric-guided-verification/
21. AI Post Transformers: MEMSEARCHER: Reinforcement Learning for LLM Memory Management — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-memsearcher-reinforcement-learning-for-l-e9ad84.mp3
22. AI Post Transformers: IMO-Bench for Robust Mathematical Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-imo-bench-for-robust-mathematical-reason-143489.mp3
23. AI Post Transformers: Breaking the Prefix Barrier with Shared KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-breaking-the-prefix-barrier-with-shared-a5e5a6.mp3
Interactive Visualization: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy