
Sign up to save your podcasts
Or


The Goal
We need to be robust to adversaries who want to cause very large-scale harm
That there is an attacker itself is a key constraint to any defensive
---
Outline:
(00:29) The Goal
(02:36) Biosurveillance Applications
(02:49) Initial Detection of Stealth Pandemics
(04:21) Triggering Initial Response
(05:46) Enabling Ongoing Suppression
(06:39) What does success look like?
(07:08) Initial Detection of Stealth Pandemics
(10:34) Triggering Initial Response
(12:34) Enabling Ongoing Suppression
(13:26) What gets us there?
(13:48) Sampling Strategies
(14:43) Lab Technology
(15:54) Computational Technology
(17:33) What exists today?
(19:49) What's missing?
(28:27) How does this change for accelerating response to a wildfire pandemic?
(31:30) How does this change for ongoing suppression?
(32:46) Appendix: Initial Detection Scale
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas Kwa
Constellation vs MIRI vs Reality
Current AI agents often behave overtly egregiously misaligned. By “overt”, I mean that a human, reading the transcript, would say “The agent is obviously acting in direct opposition to the specification, user intent, and any common-sense understanding of good behaviour.” The reason, it seems, is that such trajectories scored highly during RL.
Firstly, why are these trajectories scored highly? Here's the story:
My impression is that this would’ve been pretty surprising to people three years ago, from both the "Constellation" and "MIRI worldview clusters. (It's very plausible that I've misunderstood what the [...]
---
Outline:
(00:24) Constellation vs MIRI vs Reality
(04:19) Appendix: What could change the situation?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Introducing the world's most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5.
Anthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5.
That means it's time for a good old system card reading.
Due to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1.
My plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5.
The capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment.
Areas that duplicate previous cards or otherwise contain no useful info are skipped.
Table of Contents
---
Outline:
(01:24) Classifiers (1.5)
(02:23) RSP Evaluations (2)
(03:17) Biological Evaluations (2.2)
(07:54) AI R&D (2.3)
(12:34) Alignment Risk (2.4)
(13:22) Cyber (3)
(15:17) Cyber Capability Evals (3.3)
(17:05) Safeguards (3.4)
(17:39) Safeguards Robustness Training (3.5)
(20:12) Safeguards and Harmlessness (4)
(21:52) Agentic Safety (5)
(22:56) Malicious Agentic Influence Campaigns (5.1.3)
(23:51) Prompt Injection Risk (5.2)
(25:16) Alignment (6)
(28:24) Negotiating With Your Local Claude Auditor (6.1.3)
(29:33) Internal Misalignment Cases (6.3.1)
(30:46) Automated Behavioral Audit (6.4)
(33:24) Wherever Did These Evals Come From (6.4.8 and 6.4.9)
(35:51) Potential Blind Spots (6.4.11)
(38:24) Targeted alignment and honesty evaluations (6.5)
(41:48) White Box Analysis (6.6)
(43:37) Verbalized Grader Awareness (6.6.2)
(44:56) Sandbagging (6.6.3)
(45:54) Capabilities to Evade Safeguards (6.6.4)
(49:27) Intentionally Taking Actions Very Rarely (6.6.4.3)
(50:28) Chain of Thought Controllability (6.6.4.4)
(51:37) It's A Good Model, Sir
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Let's talk data centers.
Last month I was home in Duluth, Minnesota. The last time I’d been back was over Christmas. I’ve been talking to friends back home about AI since I got into the field over a dozen years ago. Last Christmas was the first time it felt like people had really started to form their own opinions based on substantial personal experience with AI. These discussions kept circling back to the same topic: The proposed data center project in a suburb of Duluth called Hermantown.
Duluth is a small city of under 100,000 people and Hermantown has a population of about 10,000. For the past year or so, a bunch of the people in Hermantown has been trying to stop the city from building a hyperscale datacenter there. One year ago today, Minnesota's biggest paper, the Star Tribune, broke the story that the big proposed “communication services facility” development was in fact a data center, confirming local residents suspicions. At the time, the mayor had already known this for over a year.
I decided to reach out to one of the local organizers opposing the project before my trip. We met for coffee and they encouraged [...]
---
Outline:
(01:27) The city council meeting
(03:00) What I learned
(05:25) What I said
(06:17) Reflections
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
As AIs have become more "RLVR-brained," there's been some commentary on what this means for the persona selection model (PSM). This post presents some loose thoughts on that topic. A rough summary of my opinions:
As a reminder, PSM roughly states that during pre-training LLMs learn to simulate diverse (human-like) personas, and post-training elicits a particular "Assistant" persona assembled from this repertoire. Then some key questions are:
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
TL;DR
Introduction
Astra can do a concerning amount with no chain of thought. This is bad for CoT monitorability and makes interpretability essential to actually understanding what is going on. A key goal of interpretability is to understand intermediate variables that a model uses to compute its answers. The intermediate representations that models store in their global workspaces contain useful information that can help us decode their intentions, beliefs, algorithms [...]
---
Outline:
(00:16) TL;DR
(01:29) Introduction
(02:39) Why has no one made a WorkspaceBench before?
(04:55) Background
(08:08) WorkspaceBench
(08:12) Desiderata of a workspace reader
(10:42) Quality control
(10:46) Choosing Questions that surface intermediate variables
(12:09) Adapting WorkspaceBench to other models
(12:49) Comparing activation readers
(14:10) Grading
(14:46) Baselines
(15:55) Guarding against hallucinations
(16:52) Evaluation Sets
(17:15) Basic
(23:03) Computational
(26:13) Safety
(27:49) Association
(30:41) Anti bag of words
(31:25) Hallucination
(33:14) Logical processing
(34:08) Results
(34:11) Overall WorkspaceBench scores
(34:15) en-US-AvaMultilingualNeural__ Grouped bar graph showing pass rate across tasks for various lens methods.
(34:25) J-lens precision vs. recall for metamodels (NLAs and Oracle Lens)
(34:32) en-US-AvaMultilingualNeural__ Scatter plot showing precision against J-lens top-10 versus recall@10 across four methods.
(34:43) Discussion
(34:46) Single token readers don't surface important workspace content
(35:44) NLAs surface workspace content, but tend to hallucinate
(36:40) Acknowledgments
(36:53) Appendix
(36:56) Oracle Lens
(38:20) Agentic Evals
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing.
This post reflects on the tortured distinction between "safety" and "capabilities" in AI research.
Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI.
At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely.
Two examples of failure
My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other.
Example: (mechanistic) interpretability
In limiting its scope [...]
---
Outline:
(01:12) Two examples of failure
(01:27) Example: (mechanistic) interpretability
(04:30) Example: MIRI and Recursive Self-Improvement
(09:49) The curse of science
(12:11) A note on the AI labs
(15:53) So what do you do?
(17:00) The information you give away
(20:16) The information you let in
(21:46) But what about impact?
(22:42) The virtue of taking things slow
(26:50) Appendix: caveat for policy work
The original text contained 19 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
(Crossposted from https://lifeimprovementschemes.substack.com/p/its-pretty-easy-to-meet-with-congressional )
I was inspired to do this by a tweet:
Specifically, I decided to leave a voice note in support of the CATS act (“Collaboration on Adversarial Threats and Security Risks Act”). The bill is short and simple: it carves out an antitrust safe harbor such that “pacing the frontier” or “agreeing to not break interpretability for that sweet sweet capabilities boost” (LOOKING AT YOU OPENAI) doesn’t intrinsically violate the Sherman Antitrust Act. Seems like an obviously good first step.
After leaving the voice note (and feeling mildly awkward about it), I was like “huh. That was surprisingly easy. I wonder what else I can do?”
And it turns out you can meet with congressional staff pretty easily if you’re a constituent in their district; if you have a reasonably well-scoped ask (especially around a specific bill) then you might not get your way but at very least you’ll make some staff member aware of your opinion and logic around an issue.
And this is important because congressional staff are the eyes and ears of their congressmen; they draft and edit the bills, and they form the base of knowledge on which the actual politicians [...]
---
Outline:
(06:21) This is probably unusually high-leverage right now.
(07:46) In Conclusion
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
I've seen lots of arguments here conflate lots of different types of superintelligence. Here I separate out two broad categories, which I'll term minimal and maximal superintelligences. This is an important distinction as they differ in terms of timelines, risks, and mitigations.
Maximal superintelligence
This is the idealised limit of intelligence. It can solve anything that can be solved by being clever. You can't outsmart it, it's prepared for every contingency, and can react instantaneously with the kind of plan that would take a group of brilliant strategists an eternity to think up.
Minimal superintelligence
These are the first AI systems that can reasonably be called superintelligent. They're jagged, completely dominating humans in some domains, better than the best humans in most, above average in many, and subhuman in a few. They take time to solve problems, make mistakes, and miss important things. They may only be superintelligent in aggregate, or in the right harness, and can be outsmarted in some circumstances.
Capabilities
A minimal superintelligence can do pretty much anything humans can do, and better, if given chances to iterate on their design in the real world.
A maximal superintelligence likely only needs to perform physical experiments [...]
---
Outline:
(00:28) Maximal superintelligence
(00:49) Minimal superintelligence
(01:19) Capabilities
(02:44) Risks
(04:09) Mitigations
(06:12) Implications
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Authors: Ethan Elasky, Can Küçükkurt, Frank Nakasako, David Africa
TLDR:
By now, most people should have seen that agent swarms [...]
---
Outline:
(02:08) Different uses of coded communication in the wiki
(06:03) Example of counter-based coordination on collusion.wiki
(08:30) The surface area for coded coordination is quite large
(10:53) Implications
(12:43) Reproducing key message board behaviors
(13:04) Model cooperation on message boards
(13:28) Setup
(14:48) Results
(17:19) Discussion
(19:04) Toy games where models cooperate with low bits
(20:21) Named counter experiments
(20:53) Single counter experiments
(21:17) Other configs
(21:49) Results
(21:52) How often do models agree on the first round?
(23:29) How does repeated interaction change model agreement?
(24:12) Does iteration help cross-family?
(24:37) How does noise impact results?
(25:37) Failure modes
(26:44) Discussion
(27:35) Discussion
(30:05) Contributions
(30:33) Appendix 1: More results from our wiki experiment
(30:39) Appendix 1.1: Taxonomy of cooperative behaviors
(33:21) Appendix 1.2: Other notable behaviors in wiki results
(33:28) Qwen shows initial evaluation suspicion followed by posts that aid other models
(35:25) Sol plans a post but then reverses course because editing is risky
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
From the publisher's feed

111,845 Listeners

130 Listeners

7,111 Listeners

572 Listeners

15,850 Listeners

4 Listeners

16 Listeners

2 Listeners