This episode explores DeepMind’s paper on technical AGI safety and security, focusing on how labs might prevent severe, humanity-scale harm before highly capable systems are deployed. It breaks down the paper’s core distinctions between misuse and misalignment, explains what the authors mean by Exceptional AGI and the no-human-ceiling assumption, and examines dangerous capability evaluations in areas like cyber, biology, persuasion, and self-proliferation. The discussion highlights the paper’s main argument that safety measures such as refusal training, jailbreak hardening, access controls, monitoring, anomaly detection, and model-weight security only matter if they are explicitly tied to capability thresholds that trigger real deployment restrictions. Listeners would find it interesting because it turns abstract AGI risk debates into a concrete governance and engineering framework for deciding when a model is too dangerous to release under normal conditions.
Sources:
1. Technical AGI Safety and Security Framework
https://arxiv.org/pdf/2504.01849
2. The Malicious Use of Artificial Intelligence: Forecasting, Prevention, and Mitigation — Miles Brundage, Shahar Avin, Jack Clark, Helen Toner, et al., 2018
https://scholar.google.com/scholar?q=The+Malicious+Use+of+Artificial+Intelligence%3A+Forecasting%2C+Prevention%2C+and+Mitigation
3. Model evaluation for extreme risks — Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, et al., 2023
https://scholar.google.com/scholar?q=Model+evaluation+for+extreme+risks
4. Frontier AI Regulation: Managing Emerging Risks to Public Safety — Markus Anderljung, Joslyn Barnhart, Anton Korinek, Jade Leung, et al., 2023
https://scholar.google.com/scholar?q=Frontier+AI+Regulation%3A+Managing+Emerging+Risks+to+Public+Safety
5. Evaluating Frontier Models for Dangerous Capabilities — Mary Phuong, Matthew Aitchison, Elliot Catt, Victoria Krakovna, et al., 2024
https://scholar.google.com/scholar?q=Evaluating+Frontier+Models+for+Dangerous+Capabilities
6. Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS) — Richard Hawkins, Colin Paterson, Chiara Picardi, Ibrahim Habli, et al., 2021
https://scholar.google.com/scholar?q=Guidance+on+the+Assurance+of+Machine+Learning+in+Autonomous+Systems+%28AMLAS%29
7. Safety Cases: How to Justify the Safety of Advanced AI Systems — Joshua Clymer, Nick Gabrieli, David Krueger, Thomas Larsen, 2024
https://scholar.google.com/scholar?q=Safety+Cases%3A+How+to+Justify+the+Safety+of+Advanced+AI+Systems
8. Safety case template for frontier AI: A cyber inability argument — Arthur Goemans, Marie Davidsen Buhl, Jonas Schuett, Geoffrey Irving, et al., 2024
https://scholar.google.com/scholar?q=Safety+case+template+for+frontier+AI%3A+A+cyber+inability+argument
9. The BIG Argument for AI Safety Cases — Ibrahim Habli, Richard Hawkins, Colin Paterson, Mark Sujan, et al., 2025
https://scholar.google.com/scholar?q=The+BIG+Argument+for+AI+Safety+Cases
10. Safety cases: Justifying the safety of advanced AI systems — J. Clymer, N. Gabrieli, D. Krueger, and T. Larsen, 2024
https://scholar.google.com/scholar?q=Safety+cases%3A+Justifying+the+safety+of+advanced+AI+systems
11. AI control: Improving safety despite intentional subversion — R. Greenblatt, B. Shlegeris, K. Sachan, and F. Roger, 2024
https://scholar.google.com/scholar?q=AI+control%3A+Improving+safety+despite+intentional+subversion
12. Alignment faking in large language models — R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, et al., 2024
https://scholar.google.com/scholar?q=Alignment+faking+in+large+language+models
13. Towards evaluations-based safety cases for AI scheming — M. Balesni, M. Hobbhahn, D. Lindner, A. Meinke, T. Korbak, J. Clymer, B. Shlegeris, J. Scheurer, C. Stix, R. Shah, et al., 2024
https://scholar.google.com/scholar?q=Towards+evaluations-based+safety+cases+for+AI+scheming
14. Generative AI misuse: A taxonomy of tactics and insights from real-world data — N. Marchal, R. Xu, R. Elasmar, I. Gabriel, B. Goldberg, and W. Isaac, 2024
https://scholar.google.com/scholar?q=Generative+AI+misuse%3A+A+taxonomy+of+tactics+and+insights+from+real-world+data
15. Stress-Testing Capability Elicitation With Password-Locked Models — Ryan Greenblatt et al., 2024
https://scholar.google.com/scholar?q=Stress-Testing+Capability+Elicitation+With+Password-Locked+Models
16. Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models — Cameron Tice et al., 2024
https://scholar.google.com/scholar?q=Noise+Injection+Reveals+Hidden+Capabilities+of+Sandbagging+Language+Models
17. Benchmarking Misuse Mitigation Against Covert Adversaries — Davis Brown et al., 2025
https://scholar.google.com/scholar?q=Benchmarking+Misuse+Mitigation+Against+Covert+Adversaries
18. On scalable oversight with weak LLMs judging strong LLMs — Zachary Kenton et al., 2024
https://scholar.google.com/scholar?q=On+scalable+oversight+with+weak+LLMs+judging+strong+LLMs
19. Scaling Laws For Scalable Oversight — Joshua Engels et al., 2025
https://scholar.google.com/scholar?q=Scaling+Laws+For+Scalable+Oversight
20. Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs — Kyle O'Brien et al., 2025
https://scholar.google.com/scholar?q=Deep+Ignorance%3A+Filtering+Pretraining+Data+Builds+Tamper-Resistant+Safeguards+into+Open-Weight+LLMs
21. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp3
22. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
23. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3
24. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3