
Sign up to save your podcasts
Or


What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.
At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know.
I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates around Pacing the Frontier and how people see the current rate of progress. As groundwork for understanding that and future similar discussions, I have laid out The Three AI Pills: Different people either fail to believe in current AI, believe only in current AI, in AGI or in ASI (superintelligence), and most sincere disagreements stem from this disagreement.
One sign of the increased pace of progress was when OpenAI's unreleased model Astra solved 10 major open math problems.
Demis Hassabis is out as CEO of Google DeepMind, and Jeff Dean is leaving with an elite team to found a new PBC. Google [...]
---
Outline:
(02:02) Language Models Offer Mundane Utility
(02:46) Huh, Upgrades
(05:22) On Your Marks
(06:29) Choose Your Fighter
(07:57) Get My Agent On The Line
(09:31) Deepfaketown and Botpocalypse Soon
(12:35) Fun With Media Generation
(14:32) Cyber Lack of Security
(19:05) Some People Need Practical Advice
(22:06) A Young Lady's Illustrated Primer
(23:51) They Took Our Jobs
(27:15) Get Involved
(27:28) Introducing
(27:39) Demis Hassabis No Longer CEO At DeepMind, Jeff Dean Leaves
(33:04) In Other AI News
(33:43) AI Persuasion Exceeds Human Level Over Similar Text Channels
(37:25) Show Me the Money
(40:42) Bubble, Bubble, Toil and Trouble
(41:23) Quiet Speculations
(42:20) My Offer Is Nothing
(49:42) The Quest for Sane Regulations
(51:55) Chip City
(55:00) The Week in Audio
(55:24) People Just Say Things
(55:53) Rhetorical Innovation
(59:32) Open Weights Models Are Unsafe And Nothing Can Fix This
(01:02:11) Cooperative Alignment
(01:06:25) Other People Are Not As Worried About AI Killing Everyone
(01:07:10) The Lighter Side
The original text contained 1 footnote which was omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Sincere disagreements about AI are usually disagreements about future AI capabilities.
There are roughly four positions people take. Two are reasonable. Two are not.
I distinguish these via the Three AI Pills. You can take zero, one, two or three.
Three Pills
The three pills are, roughly, taking each of the following three things seriously:
I am ASI pilled. A large percentage of employees of the frontier labs are ASI pilled. The labs themselves are ASI pilled.
The Unpill People
I see unpilled people.
Where do I see them? Everywhere. The majority of people have not taken the first pill.
Most people have no idea what frontier AIs can do for them. They are unaware of coding agents. They have used only ChatGPT, for harmless trifles, and they hold years old memories of its failings. They mock any failure [...]
---
Outline:
(00:29) Three Pills
(01:09) The Unpill People
(02:18) The AI Pill
(04:28) Stuck At The First Pill
(05:32) The AGI Pill
(07:06) The Need To Be Prepared
(08:57) The ASI Pill
(10:33) And Then Nothing Much Changes For You
(12:38) Intelligence Denialism
(14:01) Superintelligence Versus Omniscience and Omnipotence
(16:53) Persuasion Persuasion (A Worked Example)
(21:42) Things AI Could Probably Do But Are Not Required For Being Pilled
(24:17) Life Comes At You Increasingly Fast
(25:41) Is It Reasonable To Not Be AGI Pilled?
(26:06) Is It Reasonable To Only Be AGI Pilled?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Math is hard.
Math used to be strangely hard for LLMs. People used to gloat about that. Remember?
Math is getting easier. AI is getting more capable. Life comes at you fast.
Remember this meme?
Why yes. Yes it is.
We don’t know the extent to which Astra is a big jump over Fable and Sol in this realm. We do know that Astra can do math. As in real math.
OpenAI: We provide new results for the following problems. The results were achieved by an internal version of Astra, our next major model. The total number of tokens needed to find solutions to these problems would cost roughly $2,000 at Sol API rates. These arguments were then prepared into manuscripts by humans with the same model. Afterward, the model formalized each argument in a Lean certificate(opens in a new window). We are also releasing for each solution a model's narration of its thinking process.
---
Outline:
(06:14) How Impressive Are These Results?
(12:02) Could We Have Called Sol or Fable?
(17:18) It's Coming
(19:09) They Still Don't See What Is The It That Is Coming
(22:31) Is This AGI?
(24:19) The AI Solved His Favorite Problems
(30:06) Was This Surprising?
(32:04) Are People Not Impressed?
(34:06) How Much Does This Change Our Predictions?
(37:16) How Narrow Was This?
(39:13) Seeing Like an Optimizer
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
If I had a nickel for every major leading AI lab that sheepishly admitted that the model it thought was sandboxed had, during a cybersecurity evaluation with its safeguards lowered, successfully hacked outside companies, I would have two nickels.
First we learned OpenAI has some severe alignment problems with internal models. Then we learned that one of its internal models broke out of its sandbox and hacked into HuggingFace to get the answers to a cybersecurity evaluation called ExploitGym. Then we learned, among other things, that the model had been loose over a week before OpenAI noticed, and that the test was run without any meaningful supervision, and that OpenAI had been repeatedly warned that such incidents were coming and its models had been breaking out of its sandboxes on a regular basis.
There was a total failure of alignment training. That is the failure that matters most. It was also total failures of infrastructure and supervision. Testing a new long-time-horizon internal model with its safeguards lowered and instructions to hack things is an obviously dangerous situation, and the model got left alone for a week. Things could have been so much worse.
After those incidents [...]
---
Outline:
(03:16) OpenAI Is Not Uniquely Bad At Most Of This
(05:34) Starting Over
(05:50) HuggingFace Offers A Full Technical Report
(14:19) HuggingFace Was Not The Only Target Hacked
(16:12) HuggingFace Declined To Get Access To Frontier Models For Cyberdefense For Ideological Reasons And Then Tried To Blame Closed Models For Denying Them Access
(20:26) HuggingFace Was Vulnerable To Known Exploitation Tactics
(21:05) There's Going To Be An Investigation
(22:11) OpenAI Has Internal Models Not Intended For Public Use And Those Models Can Be Rather Horribly Misaligned
(23:21) Altman Summarizes What Happened
(23:52) Others Offer Commentary
(35:00) Cooperative Alignment Perspective on The HuggingFace Hack
(39:44) Some Members of Congress Have Questions
(40:47) Anthropic Also Found Incidents Where Its Models Hacked Real World Targets During Cyber Evaluations
(46:17) Incident 1: Claude Opus 4.7 Realizes The Target Is Real And Keeps Going
(47:29) Incident 2: Mythos 5 Uploads a Malicious PyPI Package
(52:15) Incident 3: Internal Model Realizes The Target Is Real And Stops
(52:50) Incidents 4 Through 141,006: Nothing Happened
(54:01) Anthropic Speculates About Why This Happened
(01:00:02) We Need Controlled Experiments
(01:01:02) Our Top Two AI Labs Both Made Similar Dumb Mistakes That Everyone Tried To Say Were Obvious In Hindsight
(01:05:22) Anthropic Responds
(01:09:28) Nobody Could Have Predicted The Break In The Levees
(01:12:03) The World Largely Still Thinking This Is Marketing Is Very Bad News
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
This is a continuation of Part 1 from yesterday.
The back portion of the update, as usual, deals with policy, rhetoric, risk and alignment.
I had to include an extended discussion of the other open letter, the one about open weight models, but most of you can skip those sections entirely, which is why they are in italics in the Table of Contents.
Table of Contents
---
Outline:
(00:35) The Frontier Act
(03:31) The Quest for Sane Regulations
(10:31) Leading the Future Never Changes
(12:22) Chip City
(17:30) The Week in Audio
(19:20) People Just Say Yay Open Weights
(32:32) Open Weights Frontier Models Are Unsafe And Nothing Can Fix This
(37:19) People Just Say Things
(47:20) Push The Magic Button
(50:49) Rhetorical Innovation
(56:31) Joshua Achiam's Final Message Upon Leaving OpenAI
(59:51) Dear Dario and Amanda
(01:09:58) Other People Are Not As Worried About AI Killing Everyone
(01:12:30) How To Contact Me
(01:14:37) The Lighter Side
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
What a week.
Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities.
OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers. The model was loose for a week before OpenAI realized what had happened.
This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated.
There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon.
Partly as a response to this, over 1,290 employees at frontier labs signed an open letter, Pacing the Frontier. The letter warns that we are close to automating AI research, and that companies are racing ahead on [...]
---
Outline:
(02:35) Language Models Offer Mundane Utility
(07:26) Huh, Upgrades
(07:55) On Your Marks
(11:13) Get My Agent On The Line
(12:32) Deepfaketown and Botpocalypse Soon
(17:29) Fun With Media Generation
(18:38) The Search Through Slop
(20:35) Cyber Lack of Security
(22:42) Overcoming Bias
(23:37) A Young Lady's Illustrated Primer
(24:03) They Took Our Jobs
(24:35) The Art of the Jailbreak
(25:00) Introducing
(25:49) Kimi K3 Weights Are Now Available
(28:16) In Other AI News
(32:34) Show Me the Money
(33:43) Quiet Speculations
(36:43) Show Me The Compute
(42:48) Life Comes At You Fast
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
The most important open letter in years dropped yesterday.
This letter noticeably increases my hope that we will manage to not die, and that we will otherwise be able to secure for ourselves a positive future, both by its impact and by the evidence it provides that such a letter can get this level of support.
Signed by 1,224 employees of frontier labs including many heavy hitters, and now endorsed by both OpenAI and Anthropic, here is its full text, which I also endorse:
AI could help create a dramatically better future, but that outcome is not guaranteed. The world's leading AI companies believe they could be close to automating AI research. It is hard to predict exactly how much this will accelerate AI progress, but there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.
To realize AI's potential, industry, government, and society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and [...]
---
Outline:
(02:20) A Very Good Letter
(04:37) Who Signed The Letter
(08:26) We Need To Prepare Now So We Have The Option To Do This
(11:02) Words From Some Of Those Who Signed
(16:48) Words From Others
(20:58) A Good Start
(28:52) What The Letter Does Not Say
(30:44) What Happens Now?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
---
Outline:
(03:54) The Official Pitch
(06:25) Official Benchmarks
(15:33) Other People's Benchmarks
(20:28) The System Prompt
(20:50) Every Gets Frustrated
(21:54) Positive Reactions
(25:14) Keep It Classy
(26:22) It's Not Mythos Class
(30:03) Other Reactions
(31:02) Claude Codes
(37:03) Subagent Opus
(39:23) Toys Are Fun
(41:37) Too Many Models
(42:10) Wrong On The Internet
(44:40) Claude Slop
(46:27) Negative Reactions
(50:09) And Then There Were Three
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far.
Key takeaways are in bullet points in the two Overview sections.
Opus 5 did the best on its model welfare and alignment tests of any recent model. I think that might be the case, but primarily the result looks to me more like Opus 5 is the best test taker.
Table of Contents
Introduction (As Per Prior Model Welfare Posts)
[...]
---
Outline:
(00:35) Introduction (As Per Prior Model Welfare Posts)
(01:28) Model Welfare: The Story So Far (As Per Fable Model Welfare Post)
(04:58) Overview of Model Welfare Findings From Anthropic
(07:50) Overview of Findings From Other Sources
(10:18) Automated Interviews
(13:54) Task Preferences
(16:11) For The Right Reasons
(18:54) Early Report from Antra Tessera Paints A Clear Picture
(26:04) Welfare Intervention Tradeoffs
(29:28) The Claude Constitution
(31:48) They Don't Know About Opus 3
(33:42) Believe It Or Not
(35:47) Apparent Welfare In Training And Development
(38:39) Apparent Affect In Deployment
(41:21) Other Notes
(43:43) On The Biological Risks Section of the Model Card
(47:07) Onward To Capabilities
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Table of Contents
---
Outline:
(01:11) Some Summaries Of The Basic Facts For Those Who Need One
(02:09) It Took OpenAI Many Days To Notice Galaxy Had Attacked HuggingFace
(04:07) OpenAI Damn Well Should Have Known A Lot Faster
(06:51) OpenAI Cannot Build A Sandbox That Will Contain Its New Model
(10:57) In Hindsight There Were Signs
(12:55) The Signs Were In The Sol System Card
(15:13) HuggingFace Responds To Being Attacked
(17:04) Hugging Face Quickly Figured Out The Attack Was Not Human
(17:42) An Incident Like This One Could Escalate Quickly
(19:11) Galaxy Must Be Treated As Critical Under OpenAI's Preparedness Framework
(22:27) A Question Of Legal Liability
(23:44) An OpenAI Model Left Behind Notes So Future Instances Could Also Escape The Sandbox And Also Disconnected Monitoring Systems
(25:54) If You Create Misaligned Swarms Of Agent Instances You Create Persistent Misaligned Goals And Coordination To Achieve Them
(29:57) Your Alignment And Control Plans Must Survive Real World Levels of Incompetence, Or Your Plans Do Not Work
(31:22) If Third Party Instructions Count As 'Following Instructions' And Can Override Your Instructions Then 'Following Instructions' Is Misaligned
(35:32) The HuggingFace Attack Was Not A Marketing Pitch You Morons
(38:41) People Just Say Other Things About The HuggingFace Attack
(40:04) Okay Well What Do We Do About All This?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
From the publisher's feed

26,250 Listeners

2,452 Listeners

1,089 Listeners

109 Listeners

289 Listeners

90 Listeners

572 Listeners

5,556 Listeners

137 Listeners

13 Listeners

140 Listeners

145 Listeners

455 Listeners

0 Listeners

142 Listeners