This episode explores whether large language models can estimate their chances of success before acting, update those estimates during a task, and use them to decide when to abstain from costly work. It explains why that matters for agentic systems in coding and software environments, where overconfidence can lead to wasted effort, unsafe actions, or expensive mistakes, and connects the paper to earlier work on calibration, uncertainty, and selective abstention. The discussion highlights the paper’s focus on three settings, including single-step coding, sequential decisions with feedback, and multi-step software engineering, while also stressing the distinction between raw capability, calibration, and rational decision-making. Listeners would find it interesting because it treats self-assessment not as a philosophical question, but as a practical requirement for building AI systems that know when not to act.
Sources:
1. Do Large Language Models Know What They Are Capable Of? — Casey O. Barkan, Sid Black, Oliver Sourbut, 2025
http://arxiv.org/abs/2512.24661
2. Language Models (Mostly) Know What They Know — Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Ethan Perez, Nicholas Joseph, and collaborators, 2022
https://scholar.google.com/scholar?q=Language+Models+%28Mostly%29+Know+What+They+Know
3. Do Large Language Models Know What They Don’t Know? — Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang, 2023
https://scholar.google.com/scholar?q=Do+Large+Language+Models+Know+What+They+Don%E2%80%99t+Know%3F
4. Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception — Shiyu Ni, Keping Bi, Jiafeng Guo, Lulu Yu, Baolong Bi, Xueqi Cheng, 2025
https://scholar.google.com/scholar?q=Towards+Fully+Exploiting+LLM+Internal+States+to+Enhance+Knowledge+Boundary+Perception
5. What Large Language Models Know and What People Think They Know — Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Lukas W. Mayer, Padhraic Smyth, and collaborators, 2025
https://scholar.google.com/scholar?q=What+Large+Language+Models+Know+and+What+People+Think+They+Know
6. Teaching Models to Express Their Uncertainty in Words — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
https://scholar.google.com/scholar?q=Teaching+Models+to+Express+Their+Uncertainty+in+Words
7. A Survey of Confidence Estimation and Calibration in Large Language Models — Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, Iryna Gurevych, 2024
https://scholar.google.com/scholar?q=A+Survey+of+Confidence+Estimation+and+Calibration+in+Large+Language+Models
8. Calibrating Language Models with Adaptive Temperature Scaling — Johnathan Xie, Annie S. Chen, Yoonho Lee, Eric Mitchell, Chelsea Finn, 2024
https://scholar.google.com/scholar?q=Calibrating+Language+Models+with+Adaptive+Temperature+Scaling
9. Calibrating the Confidence of Large Language Models by Eliciting Fidelity — Mozhi Zhang, Mianqiu Huang, Rundong Shi, Linsen Guo, Chong Peng, Peng Yan, Yaqian Zhou, Xipeng Qiu, 2024
https://scholar.google.com/scholar?q=Calibrating+the+Confidence+of+Large+Language+Models+by+Eliciting+Fidelity
10. Selective Classification for Deep Neural Networks — Yonatan Geifman, Ran El-Yaniv, 2017
https://scholar.google.com/scholar?q=Selective+Classification+for+Deep+Neural+Networks
11. SelectiveNet: A Deep Neural Network with an Integrated Reject Option — Yonatan Geifman, Ran El-Yaniv, 2019
https://scholar.google.com/scholar?q=SelectiveNet%3A+A+Deep+Neural+Network+with+an+Integrated+Reject+Option
12. Selective-LAMA: Selective Prediction for Confidence-Aware Evaluation of Language Models — Hiyori Yoshikawa, Naoaki Okazaki, 2023
https://scholar.google.com/scholar?q=Selective-LAMA%3A+Selective+Prediction+for+Confidence-Aware+Evaluation+of+Language+Models
13. Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference — Bo-Wei Chen, Chung-Chi Chen, An-Zi Yen, 2026
https://scholar.google.com/scholar?q=Confidence-Driven+Multi-Scale+Model+Selection+for+Cost-Efficient+Inference
14. Quantifying Uncert-AI-nty: Testing the Accuracy of LLMs' Confidence Judgments — Trent N. Cash, Daniel M. Oppenheimer, Sara Christie, Mira Devgan, 2025
https://scholar.google.com/scholar?q=Quantifying+Uncert-AI-nty%3A+Testing+the+Accuracy+of+LLMs%27+Confidence+Judgments
15. Credence Calibration Game? Calibrating Large Language Models Through Structured Play — Ke Fang, Tianyi Zhao, Lu Cheng, 2025
https://scholar.google.com/scholar?q=Credence+Calibration+Game%3F+Calibrating+Large+Language+Models+Through+Structured+Play
16. Calibration and Correctness of Language Models for Code — Claudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel, Md Rafiqul Islam Rabin, Amin Alipour, Susmit Jha, Prem Devanbu, Toufique Ahmed, 2025
https://scholar.google.com/scholar?q=Calibration+and+Correctness+of+Language+Models+for+Code
17. Large Language Models Must Be Taught to Know What They Don't Know — Sanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins, Arka Pal, Umang Bhatt, Adrian Weller, Samuel Dooley, Micah Goldblum, Andrew G. Wilson, 2024
https://scholar.google.com/scholar?q=Large+Language+Models+Must+Be+Taught+to+Know+What+They+Don%27t+Know
18. SMARTCAL: An Approach to Self-Aware Tool-Use Evaluation and Calibration — Yuanhao Shen, Xiaodan Zhu, Lei Chen, 2024
https://scholar.google.com/scholar?q=SMARTCAL%3A+An+Approach+to+Self-Aware+Tool-Use+Evaluation+and+Calibration
19. Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations — Yuxi Xia, Dennis Ulmer, Terra Blevins, Yihong Liu, Hinrich Schutze, Benjamin Roth, 2026
https://scholar.google.com/scholar?q=Calibration+Is+Not+Enough%3A+Evaluating+Confidence+Estimation+Under+Language+Variations
20. CoT-UQ: Improving Response-wise Uncertainty Quantification in LLMs with Chain-of-Thought — Boxuan Zhang, Ruqi Zhang, 2025
https://scholar.google.com/scholar?q=CoT-UQ%3A+Improving+Response-wise+Uncertainty+Quantification+in+LLMs+with+Chain-of-Thought
21. Structured Uncertainty guided Clarification for LLM Agents — Manan Suri, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan A. Rossi, Dinesh Manocha, 2025
https://scholar.google.com/scholar?q=Structured+Uncertainty+guided+Clarification+for+LLM+Agents
22. CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty — Johannes Kirmayr, Lukas Stappen, Elisabeth Andre, 2026
https://scholar.google.com/scholar?q=CAR-bench%3A+Evaluating+the+Consistency+and+Limit-Awareness+of+LLM+Agents+under+Real-World+Uncertainty
23. Improving Interactive In-Context Learning from Natural Language Feedback — Martin Klissarov, Jonathan Cook, Diego Antognini, Hao Sun, Jingling Li, Natasha Jaques, Claudiu Musat, Edward Grefenstette, 2026
https://scholar.google.com/scholar?q=Improving+Interactive+In-Context+Learning+from+Natural+Language+Feedback
24. AI Post Transformers: Benchmarking Test-Time Scaling for General LLM Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-benchmarking-test-time-scaling-for-gener-8f14f9.mp3
25. AI Post Transformers: Program Synthesis with Large Language Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-program-synthesis-with-large-language-mo-b962ec.mp3
26. AI Post Transformers: SkillsBench for Evaluating Agent Skills — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-skillsbench-for-evaluating-agent-skills-58bb1e.mp3
27. AI Post Transformers: Stabilizing Efficient Reasoning with Step-Level Advantage Selection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-29-stabilizing-efficient-reasoning-with-ste-1e589d.mp3
28. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
Interactive Visualization: Can LLMs Judge Their Own Capabilities?