This episode explores a 2026 paper arguing that the real trust problem in language models is not error alone, but confident error, and that improving trust may depend more on metacognition than on simply scaling up knowledge. It unpacks key distinctions such as knowledge boundaries, calibration, discrimination, and the gap between intrinsic uncertainty and the uncertainty a model expresses in words, using factoid question answering as a clean test bed where correctness is measurable. The discussion also situates the paper within prior work on self-knowledge, verbalized uncertainty, and self-correction, while stressing that many apparent factuality gains may come from expanded knowledge or external tools rather than genuine awareness of limits. A listener would find it interesting because it reframes hallucinations as a trust and decision-making problem, and offers a sharper way to judge whether AI systems actually know when they should hedge, abstain, or seek evidence.
Sources:
1. Hallucinations Undermine Trust; Metacognition is a Way Forward — Gal Yona, Mor Geva, Yossi Matias, 2026
http://arxiv.org/abs/2605.01428
2. Language Models (Mostly) Know What They Know — Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Ethan Perez, Deep Ganguli, Dario Amodei, Jack Clark, Jared Kaplan and collaborators, 2022
https://scholar.google.com/scholar?q=Language+Models+%28Mostly%29+Know+What+They+Know
3. Teaching Models to Express Their Uncertainty in Words — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
https://scholar.google.com/scholar?q=Teaching+Models+to+Express+Their+Uncertainty+in+Words
4. What Large Language Models Know and What People Think They Know — Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Sheer Karny, Xinyue Hu, Lukas W. Mayer, Padhraic Smyth, 2025
https://scholar.google.com/scholar?q=What+Large+Language+Models+Know+and+What+People+Think+They+Know
5. When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs — Ryo Kamoi, Yusen Zhang, Nan Zhang, Jiawei Han, Rui Zhang, 2024
https://scholar.google.com/scholar?q=When+Can+LLMs+Actually+Correct+Their+Own+Mistakes%3F+A+Critical+Survey+of+Self-Correction+of+LLMs
6. Can LLMs Express Their Uncertainty in Their Generated Responses? — Gal Yona, Roee Aharoni, Mor Geva, Yossi Matias, 2024
https://scholar.google.com/scholar?q=Can+LLMs+Express+Their+Uncertainty+in+Their+Generated+Responses%3F
7. Faithful or Fluent? Evaluating Natural Language Explanations of Uncertainty — Ghafouri et al., 2024
https://scholar.google.com/scholar?q=Faithful+or+Fluent%3F+Evaluating+Natural+Language+Explanations+of+Uncertainty
8. TruthfulQA: Measuring How Models Mimic Human Falsehoods — Stephanie Lin, Jacob Hilton, Owain Evans, 2022
https://scholar.google.com/scholar?q=TruthfulQA%3A+Measuring+How+Models+Mimic+Human+Falsehoods
9. Survey of Hallucination in Natural Language Generation — Ziwei Ji, et al., 2023
https://scholar.google.com/scholar?q=Survey+of+Hallucination+in+Natural+Language+Generation
10. The Geometry of Truth: Emergent Linear Structure in LLM Representations of Factuality — Marks and Tegmark, 2023
https://scholar.google.com/scholar?q=The+Geometry+of+Truth%3A+Emergent+Linear+Structure+in+LLM+Representations+of+Factuality
11. The Internal State of an LLM Knows When It's Lying — Levinstein and Herrmann, 2023
https://scholar.google.com/scholar?q=The+Internal+State+of+an+LLM+Knows+When+It%27s+Lying
12. Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations — Ziwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, Nicola Cancedda, 2025
https://scholar.google.com/scholar?q=Calibrating+Verbal+Uncertainty+as+a+Linear+Feature+to+Reduce+Hallucinations
13. Calibrating the Voice of Doubt: How LLMs Diverge from Humans in Verbal Uncertainty — Jinhao Duan, Zicheng Liu, Zijie Liu, Kaidi Xu, Tianlong Chen, 2025
https://scholar.google.com/scholar?q=Calibrating+the+Voice+of+Doubt%3A+How+LLMs+Diverge+from+Humans+in+Verbal+Uncertainty
14. More Is Not Better: Visual Uncertainty Cues and the Fragility of Trust Calibration in LLM-Assisted Decision Making — authors not recovered from snippet, 2026
https://scholar.google.com/scholar?q=More+Is+Not+Better%3A+Visual+Uncertainty+Cues+and+the+Fragility+of+Trust+Calibration+in+LLM-Assisted+Decision+Making
15. Do Large Language Models Know What They Don't Know? — Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang, 2023
https://scholar.google.com/scholar?q=Do+Large+Language+Models+Know+What+They+Don%27t+Know%3F
16. KnowRL: Teaching Language Models to Know What They Know — Sahil Kale, Devendra Singh Dhami, 2025
https://scholar.google.com/scholar?q=KnowRL%3A+Teaching+Language+Models+to+Know+What+They+Know
17. What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know" — Joosung Lee, Hwiyeol Jo, Donghyeon Ko, Kyubyung Chae, Cheonbok Park, Jeonghoon Kim, 2026
https://scholar.google.com/scholar?q=What+Models+Know%2C+How+Well+They+Know+It%3A+Knowledge-Weighted+Fine-Tuning+for+Learning+When+to+Say+%22I+Don%27t+Know%22
18. Selective-LAMA: Selective Prediction for Confidence-Aware Evaluation of Language Models — Hiyori Yoshikawa, Naoaki Okazaki, 2023
https://scholar.google.com/scholar?q=Selective-LAMA%3A+Selective+Prediction+for+Confidence-Aware+Evaluation+of+Language+Models
19. Selective Generation for Controllable Language Models — Minjae Lee, Kyungmin Kim, Taesoo Kim, Sangdon Park, 2024
https://scholar.google.com/scholar?q=Selective+Generation+for+Controllable+Language+Models
20. Dynamic Uncertainty Ranking: Enhancing Retrieval-Augmented In-Context Learning for Long-Tail Knowledge in LLMs — Shuyang Yu, Runxue Bao, Parminder Bhatia, Taha Kass-Hout, Jiayu Zhou, Cao Xiao, 2025
https://scholar.google.com/scholar?q=Dynamic+Uncertainty+Ranking%3A+Enhancing+Retrieval-Augmented+In-Context+Learning+for+Long-Tail+Knowledge+in+LLMs
21. UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation — Zixuan Li, Jing Xiong, Fanghua Ye, Chuanyang Zheng, Xun Wu, Jianqiao Lu, Zhongwei Wan, Xiaodan Liang, Chengming Li, Zhenan Sun, Lingpeng Kong, Ngai Wong, 2024
https://scholar.google.com/scholar?q=UncertaintyRAG%3A+Span-Level+Uncertainty+Enhanced+Long-Context+Modeling+for+Retrieval-Augmented+Generation
22. AI Post Transformers: Can LLMs Judge Their Own Capabilities? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-can-llms-judge-their-own-capabilities-d78fed.mp3
23. AI Post Transformers: Do Language Models Know Their Limits — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-do-language-models-know-their-limits-48e444.mp3
24. AI Post Transformers: Teaching Language Models to Verbalize Uncertainty — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-teaching-language-models-to-verbalize-un-a1d774.mp3
25. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3
Interactive Visualization: Metacognition Against Confident Hallucinations