The Don’t Worry About the Vase Podcast is a listener-supported podcast. To receive new posts and support the cost of creation, consider becoming a free or paid subscriber.
This has been an Askwho Casts audio conversion. If you would like your own private feed of audio conversions of any blog posts you would like to listen to, You can sign up For Askwho Casts Pro at https://app.askwhocasts.com/, Where you can give any post the multi-voiced podcast treatment, into your own podcast feed. Thanks for listening.
* 00:00:00 - Introduction
* 00:01:07 - Agent Model 1 and Agent Model 2
* 00:02:53 - Executive Summary (1)
* 00:04:17 - The Rules Are Serious But Not Literal
* 00:06:39 - Misalignment Is a State of Mind (2.5)
* 00:12:28 - Autonomy Threat Model 1: Misalignment in High-Stakes Settings (2)
* 00:17:30 - Some Strange Uses Of The Word Safe I Wasn’t Previously Aware Of
* 00:19:19 - Now Versus Future (2.17)
* 00:19:55 - The Core Claims And Argument (2.6)
* 00:31:11 - The Rest of the Important Arguments In Section 2
* 00:33:48 - Risk Assessment (2.19)
* 00:34:41 - Pre-Internal-Deployment Review (2.18)
* 00:35:44 - A Guide To Internal Use Monitoring (2.23.1)
* 00:43:06 - Blocking Interventions (2.23.2)
* 00:44:25 - The Power Seeking Environment Evaluation (2.24)
* 00:45:27 - Opus 4.8-Reward-Hacker (2.25)
* 00:47:34 - Autonomy threat model 2: Risks from automated R&D (3)
* 00:51:03 - Yes That Does Seem Kind Of Risky
* 00:53:41 - Could We Replace Our Researchers?
* 00:55:30 - How Much Could We Be Accelerating Our AI Researchers?
* 00:58:53 - What Could Possibly Go Wrong If We Replaced Our Researchers?
* 00:59:35 - Risk Mitigations For AI R&D Automation
* 01:08:08 - Overall Risk From Automation of AI R&D
* 01:08:22 - Biological and Technically Also Chemical Weapons Production
* 01:12:05 - The Threat Models for Biological and Chemical Weapons
* 01:17:35 - Model Capabilities (4.4)
* 01:19:06 - Classifiers (4.5)
* 01:21:07 - Acceleration Dynamics (5.1)
* 01:22:09 - Distillation (5.1.1)
* 01:23:22 - Safety Process Failures (5.2)
* 01:23:34 - Refusing To Find Innovative Misalignment Techniques (5.2.2)
* 01:24:55 - Exposing the Chain of Thought Reasoning To Grading Pressure Quite a Lot (5.2.3)
* 01:26:27 - Directly Training On Misaligned Behavior During a Production Training Run (5.2.4)
* 01:29:06 - An instance of unmonitored unrestricted agents with access to sensitive resources (5.2.5)
* 01:30:06 - Repeated training on alignment-faking transcript datasets (5.2.6)
* 01:32:47 - Benefits From Anthropic’s Operating as a Frontier AI company (5.3)
* 01:34:57 - Model Weight Security (6.4)
* 01:35:17 - Risk Has Been Reported
https://thezvi.substack.com/p/anthropic-risk-report-august-2026?r=67y1h&utm_campaign=post-expanded-share&utm_medium=web
Get full access to DWAtV Podcast at dwatvpodcast.substack.com/subscribe