This episode explores ReasonCACHE, a method for improving multi-step reasoning in large language models by keeping the backbone frozen and training a compact per-layer key-value memory instead of updating billions of weights. It situates the paper against in-context learning, many-shot prompting, prefix tuning, LoRA, and context-distillation work, explaining how learned latent memory sits between raw prompting and full fine-tuning. The discussion centers on the paper’s real claim and its main point of skepticism: whether these learned caches actually teach a reusable reasoning procedure or mostly compress and elicit abilities the model already had. Listeners would find it interesting because it connects a concrete new method to a larger debate about how LLMs acquire reasoning skills, while also highlighting the practical payoff of avoiding huge prompts, quadratic attention costs, and brittle long-context setups.
Sources:
1. ReasonCACHE: Teaching LLMs To Reason Without Weight Updates — Sharut Gupta, Phillip Isola, Stefanie Jegelka, David Lopez-Paz, Kartik Ahuja, Mark Ibrahim, Mohammad Pezeshki, 2026
http://arxiv.org/abs/2602.02366
2. Prefix-Tuning: Optimizing Continuous Prompts for Generation — Xiang Lisa Li, Percy Liang, 2021
https://arxiv.org/abs/2101.00190
3. The Power of Scale for Parameter-Efficient Prompt Tuning — Brian Lester, Rami Al-Rfou, Noah Constant, 2021
https://arxiv.org/abs/2104.08691
4. P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks — Xiao Liu, Kaixuan Ji, Yicheng Fu, Zhengxiao Du, Zhilin Yang, Jie Tang, 2022
https://arxiv.org/abs/2110.07602
5. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Weizhu Chen, et al., 2021
https://arxiv.org/abs/2106.09685
6. Adapting Language Models to Compress Contexts — Alexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi Chen, 2023
https://arxiv.org/abs/2305.14788
7. Learning to Compress Prompts with Gist Tokens — Jesse Mu, Xiang Lisa Li, Noah Goodman, 2023
https://arxiv.org/abs/2304.08467
8. Deliberation in Latent Space via Differentiable Cache Augmentation — Luyang Liu, Jonas Pfeiffer, Jiaxing Wu, Jun Xie, Arthur Szlam, 2024
https://arxiv.org/abs/2412.17747
9. When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations — Aleksandar Petrov, Philip H. S. Torr, Adel Bibi, 2023
https://scholar.google.com/scholar?q=When+Do+Prompting+and+Prefix-Tuning+Work%3F+A+Theory+of+Capabilities+and+Limitations
10. Many-Shot In-Context Learning — Rishabh Agarwal et al., 2024
https://scholar.google.com/scholar?q=Many-Shot+In-Context+Learning
11. Cartridges: Lightweight and general-purpose long context representations via self-study — Sabri Eyuboglu et al., 2025
https://scholar.google.com/scholar?q=Cartridges%3A+Lightweight+and+general-purpose+long+context+representations+via+self-study
12. Great Memory, Shallow Reasoning: Limits of kNN-LMs — Shangyi Geng, Wenting Zhao, Alexander M. Rush, 2024
https://scholar.google.com/scholar?q=Great+Memory%2C+Shallow+Reasoning%3A+Limits+of+kNN-LMs
13. Training Plug-n-Play Knowledge Modules with Deep Context Distillation — Lucas Caccia, Alan Ansell, Edoardo Ponti, Ivan Vulić, Alessandro Sordoni, 2025
https://scholar.google.com/scholar?q=Training+Plug-n-Play+Knowledge+Modules+with+Deep+Context+Distillation
14. More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives — Xiaoqing Zhang et al., 2025
https://scholar.google.com/scholar?q=More+is+not+always+better%3F+Enhancing+Many-Shot+In-Context+Learning+with+Differentiated+and+Reweighting+Objectives
15. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han, Rajesh Jayaram, Amin Karbasi, Vahab Mirrokni, David P. Woodruff, Amir Zandieh, 2023
https://scholar.google.com/scholar?q=HyperAttention%3A+Long-context+Attention+in+Near-Linear+Time
16. Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning — Ling Team et al., 2025
https://scholar.google.com/scholar?q=Every+Attention+Matters%3A+An+Efficient+Hybrid+Architecture+for+Long-Context+Reasoning
17. AI Post Transformers: When Many-Shot CoT Becomes Test-Time Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-when-many-shot-cot-becomes-test-time-lea-c25bfe.mp3
18. AI Post Transformers: Can Models Learn from Long Context? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-can-models-learn-from-long-context-77533e.mp3
19. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
20. AI Post Transformers: Latent Reasoning with Normalizing Flows — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-06-latent-reasoning-with-normalizing-flows-6ee916.mp3
21. AI Post Transformers: Training LLMs for Divide-and-Conquer Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-training-llms-for-divide-and-conquer-rea-ea6e22.mp3
22. AI Post Transformers: Why Open Relational Foundation Models Fail — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-22-why-open-relational-foundation-models-fa-c303c6.mp3
Interactive Visualization: ReasonCACHE: Learning Reasoning Without Weight Updates