AI Post Transformers

When Many-Shot CoT Becomes Test-Time Learning


Listen Later

This episode explores a May 13, 2026 arXiv paper arguing that many-shot chain-of-thought prompting can act less like simple retrieval and more like a form of test-time learning, where structured context helps a model reason during inference without changing its weights. It examines the paper’s main claims that adding many reasoning demonstrations does not reliably help across all settings, that semantically similar examples can fail when their reasoning procedures are not actually usable, and that the order of demonstrations becomes more important as prompts grow longer. The discussion also focuses on the paper’s Curvilinear Demonstration Selection approach, which treats prompt construction more like designing a lesson plan than doing nearest-neighbor search, with especially notable gains on geometry tasks. Listeners would find it interesting because it challenges a common assumption behind huge context windows: more examples are not automatically better, and effective prompting may depend on curriculum design, model capabilities, and the compatibility of reasoning traces.
Sources:
1. When Many-Shot CoT Becomes Test-Time Learning
https://arxiv.org/pdf/2605.13511
2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, Denny Zhou, 2022
https://arxiv.org/abs/2201.11903
3. Large Language Models are Zero-Shot Reasoners — Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, Yusuke Iwasawa, 2022
https://arxiv.org/abs/2205.11916
4. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, Denny Zhou, 2022
https://arxiv.org/abs/2203.11171
5. Automatic Chain of Thought Prompting in Large Language Models — Zhuosheng Zhang, Aston Zhang, Mu Li, Alex Smola, 2022
https://arxiv.org/abs/2210.03493
6. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2020
https://arxiv.org/abs/1909.13231
7. What Learning Algorithm is In-Context Learning? Investigations with Linear Models — Ekin Akyurek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, Denny Zhou, 2022
https://arxiv.org/abs/2211.15661
8. Transformers Learn In-Context by Gradient Descent — Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, Joao Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, Max Vladymyrov, 2022
https://arxiv.org/abs/2212.07677
9. Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters — Charlie Snell, Jaehoon Lee, Kelvin Xu, Aviral Kumar, 2024
https://arxiv.org/abs/2408.03314
10. What Makes Good In-Context Examples for GPT-3? — Jiachang Liu, Dinghan Shen, Yizhe Zhang, Bill Dolan, Lawrence Carin, Weizhu Chen, 2021
https://arxiv.org/abs/2101.06804
11. Learning To Retrieve Prompts for In-Context Learning — Ohad Rubin, Jonathan Herzig, Jonathan Berant, 2021
https://arxiv.org/abs/2112.08633
12. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity — Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, Pontus Stenetorp, 2021
https://arxiv.org/abs/2104.08786
13. Many-Shot CoT-ICL: Making In-Context Learning Truly Learn — Tsz Ting Chung, Lemao Liu, Mo Yu, Dit-Yan Yeung, 2026
https://arxiv.org/abs/2605.13511
14. Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? — Sewon Min, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer, 2022
https://scholar.google.com/scholar?q=Rethinking+the+Role+of+Demonstrations%3A+What+Makes+In-Context+Learning+Work%3F
15. Test-Time Compute: Scaling Language Models with More Thinking — OpenAI, 2024
https://scholar.google.com/scholar?q=Test-Time+Compute%3A+Scaling+Language+Models+with+More+Thinking
16. ALR2: A Retrieve-then-Reason Framework for Long-context Question Answering — Huayang Li, Pat Verga, Priyanka Sen, Bowen Yang, Vijay Viswanathan, Patrick Lewis, Taro Watanabe, Yixuan Su, 2024
https://scholar.google.com/scholar?q=ALR2%3A+A+Retrieve-then-Reason+Framework+for+Long-context+Question+Answering
17. Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models — Yifu Qiu, Varun Embar, Yizhe Zhang, Navdeep Jaitly, Shay B. Cohen, Benjamin Han, 2025
https://scholar.google.com/scholar?q=Eliciting+In-context+Retrieval+and+Reasoning+for+Long-context+Large+Language+Models
18. Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions — Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish Sabharwal, 2023
https://scholar.google.com/scholar?q=Interleaving+Retrieval+with+Chain-of-Thought+Reasoning+for+Knowledge-Intensive+Multi-Step+Questions
19. Active Prompting with Chain-of-Thought for Large Language Models — Shizhe Diao, Pengcheng Wang, Yong Lin, Rui Pan, Xiang Liu, Tong Zhang, 2023
https://scholar.google.com/scholar?q=Active+Prompting+with+Chain-of-Thought+for+Large+Language+Models
20. Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models — Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, Weizhu Chen, 2023
https://scholar.google.com/scholar?q=Synthetic+Prompting%3A+Generating+Chain-of-Thought+Demonstrations+for+Large+Language+Models
21. Contrastive Chain-of-Thought Prompting — Yew Ken Chia, Guizhen Chen, Luu Anh Tuan, Soujanya Poria, Lidong Bing, 2023
https://scholar.google.com/scholar?q=Contrastive+Chain-of-Thought+Prompting
22. Let's (not) just put things in Context: Test-Time Training for Long-Context LLMs — Rachit Bansal, Aston Zhang, Rishabh Tiwari, Lovish Madaan, Sai Surya Duvvuri, Devvrit Khatri, David Brandfonbrener, David Alvarez-Melis, Prajjwal Bhargava, Mihir Sanjay Kale, Samy Jelassi, 2025
https://scholar.google.com/scholar?q=Let%27s+%28not%29+just+put+things+in+Context%3A+Test-Time+Training+for+Long-Context+LLMs
23. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
24. AI Post Transformers: Training LLMs for Divide-and-Conquer Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-training-llms-for-divide-and-conquer-rea-ea6e22.mp3
25. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
26. AI Post Transformers: Test-time Scaling for Multi-Agent Collaborative Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-test-time-scaling-for-multi-agent-collab-082570.mp3
27. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
Interactive Visualization: When Many-Shot CoT Becomes Test-Time Learning
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof