AI Post Transformers

Do Language Models Know Their Limits


Listen Later

This episode explores whether large language models can genuinely recognize the limits of their own knowledge or whether they have simply learned to sound uncertain in socially acceptable ways. It examines the paper’s idea of “self-knowledge” through the lens of confidence calibration, including the dangerous case where a model does not know an answer but responds with unwarranted confidence. The discussion walks through the SelfAware benchmark, explaining how it pairs unanswerable questions with semantically similar answerable ones and why that design is both insightful and methodologically slippery. Listeners would find it interesting because it gets past simple accuracy scores and asks a more consequential question for AI safety and product reliability: when a model says “I don’t know,” is that real judgment or just polished behavior?
Sources:
1. Do Large Language Models Know What They Don't Know? — Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang, 2023
http://arxiv.org/abs/2305.18153
2. Finetuned Language Models Are Zero-Shot Learners — Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew Dai, Quoc V. Le, 2021
https://scholar.google.com/scholar?q=Finetuned+Language+Models+Are+Zero-Shot+Learners
3. Multitask Prompted Training Enables Zero-Shot Task Generalization — Victor Sanh, Albert Webson, Colin Raffel and many coauthors, 2021
https://scholar.google.com/scholar?q=Multitask+Prompted+Training+Enables+Zero-Shot+Task+Generalization
4. Training Language Models to Follow Instructions with Human Feedback — Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida and many coauthors, 2022
https://scholar.google.com/scholar?q=Training+Language+Models+to+Follow+Instructions+with+Human+Feedback
5. Scaling Instruction-Finetuned Language Models — Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, Xuezhi Wang, Denny Zhou, Quoc V. Le, Jason Wei and many coauthors, 2022
https://scholar.google.com/scholar?q=Scaling+Instruction-Finetuned+Language+Models
6. Self-Instruct: Aligning Language Models with Self-Generated Instructions — Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi, 2023
https://scholar.google.com/scholar?q=Self-Instruct%3A+Aligning+Language+Models+with+Self-Generated+Instructions
7. Measuring and Improving Factuality in Large Language Models with Calibrated Confidence Scores — Saurav Kadavath, Eric Wallace, Luyu Gao, et al., 2022
https://scholar.google.com/scholar?q=Measuring+and+Improving+Factuality+in+Large+Language+Models+with+Calibrated+Confidence+Scores
8. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models — Aarohi Srivastava, Jos Rozen, Francesco Tintarev, et al., 2022
https://scholar.google.com/scholar?q=Beyond+the+Imitation+Game%3A+Quantifying+and+extrapolating+the+capabilities+of+language+models
9. Teaching Small Language Models to Reason — Jason Wei, Xuezhi Wang, Dale Schuurmans, et al., 2022
https://scholar.google.com/scholar?q=Teaching+Small+Language+Models+to+Reason
10. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, et al., 2022
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
11. SimCSE: Simple Contrastive Learning of Sentence Embeddings — Tianyu Gao, Xingcheng Yao, Danqi Chen, 2021
https://scholar.google.com/scholar?q=SimCSE%3A+Simple+Contrastive+Learning+of+Sentence+Embeddings
12. SQuAD 2.0: The Stanford Question Answering Dataset — Pranav Rajpurkar, Robin Jia, Percy Liang, 2018
https://scholar.google.com/scholar?q=SQuAD+2.0%3A+The+Stanford+Question+Answering+Dataset
13. Uncertainty Distillation: Teaching Language Models to Express Semantic Confidence — Sophia Hager et al., 2025
https://scholar.google.com/scholar?q=Uncertainty+Distillation%3A+Teaching+Language+Models+to+Express+Semantic+Confidence
14. Large Language Model Uncertainty Measurement and Calibration for Medical Diagnosis and Treatment — Thomas Savage et al., 2024
https://scholar.google.com/scholar?q=Large+Language+Model+Uncertainty+Measurement+and+Calibration+for+Medical+Diagnosis+and+Treatment
15. Uncertainty Quantification and Confidence Calibration in Large Language Models: A Survey — Xiaoou Liu et al., 2025
https://scholar.google.com/scholar?q=Uncertainty+Quantification+and+Confidence+Calibration+in+Large+Language+Models%3A+A+Survey
16. Unanswerability Evaluation for Retrieval Augmented Generation — Xiangyu Peng, Prafulla Kumar Choubey, Caiming Xiong, Chien-Sheng Wu, 2025
https://scholar.google.com/scholar?q=Unanswerability+Evaluation+for+Retrieval+Augmented+Generation
17. Answerability in Retrieval-Augmented Open-Domain Question Answering — Rustam Abdumalikov, Pasquale Minervini, Yova Kementchedjhieva, 2024
https://scholar.google.com/scholar?q=Answerability+in+Retrieval-Augmented+Open-Domain+Question+Answering
18. The Art of Saying No: Contextual Noncompliance in Language Models — Faeze Brahman et al., 2024
https://scholar.google.com/scholar?q=The+Art+of+Saying+No%3A+Contextual+Noncompliance+in+Language+Models
19. AI Post Transformers: Hallucination to Truth: A Review of Fact-Checking and Factuality Evaluation in Large Language Model — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/hallucination-to-truth-a-review-of-fact-checking-and-factuality-evaluation-in-la/
20. AI Post Transformers: Internal Safety Collapse in Frontier LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-internal-safety-collapse-in-frontier-llm-8be72f.mp3
21. AI Post Transformers: Self-Search Reinforcement Learning for LLMs — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/self-search-reinforcement-learning-for-llms/
22. AI Post Transformers: Real Context Size and Context Rot — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-real-context-size-and-context-rot-56cbb4.mp3
Interactive Visualization: Do Language Models Know Their Limits
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof