
Sign up to save your podcasts
Or


Oh, good. They noticed.
Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval.
Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally.
As in, Anthropic paused its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act.
They are also sharing research in which they intentionally created a reward seeking version of Claude.
Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable.
Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post.
Table of Contents
---
Outline:
(01:16) This Just In
(02:43) Anthropic Parallel Pauses
(08:22) Pause The Data Brokers
(09:54) Pacing the Frontier
(11:54) Misalignment Assessment
(13:39) Defects In Training Environments Disproportionately Cause Cheating
(14:59) Creating Reward Hacker Opus
(19:33) Undo It
(21:00) Mistakes Were Made
(23:33) Internal Security Posture
(25:26) One Does Not Simply Fix The RL Environments
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Okay, so we who read blogs like this one have collectively realized there really is a lot going on right now. There is Big Trouble in Baby Superintelligence.
So how do we get the rest of the world to take it appropriately seriously? Where do we go from here? Not only what can we do to not have a worse version of this happen again, but to ensure good outcomes generally, and employ what we learned?
There are a lot of ideas out there. OpenAI is going to be implementing some of them, at substantial cost, since the cost of not doing so is clearly far higher, even short term. My worry continues to be that their fundamental approach is fatally flawed, and they are not focusing on the right things.
It is highly fortunate that the OpenAI agents hacked HuggingFace. This is the only reason we know about all the severe internal failures at OpenAI, and gives us an opportunity to wake up before it is too late.
We do not have enough details to know what happened internally, both before and after the attack, and might never know. Before the attack, various internal [...]
---
Outline:
(03:35) Nothing Matters, Says Mainstream Media
(06:27) Move Along, Nothing To See Here
(12:40) Do They Realize They Are Not The Good Guys?
(17:22) Very Serious People
(31:30) What's In a Name?
(34:05) Learn Neuralese In Three Easy Steps
(35:37) Dwarkesh Patel Realizes He Ran A Natural Experiment
(40:40) Politicians Take Notice
(44:47) Pick Up The Phone
(46:40) A Failure To Communicate
(49:00) Anthony Aguirre Goes Over What We Learned
(50:28) Trying To Solve The Wrong Problems Using The Wrong Methods Based On A Wrong Model Of The World Derived From Poor Thinking And Hoping All Of Your Mistakes Will Cancel Out
(55:28) Indirect Pressure on the Chain of Thought
(56:39) A Matter of Trust
(59:21) Blowing the Whistle
(01:04:40) The Punishment For Being Late Is Death
(01:12:52) Another Kind Of Law
(01:16:13) What Is The Law?
(01:17:46) Building On Success
(01:19:49) Total Research Transparency
(01:21:20) Yo Shavit Calls For Widespread Disclosure Of Misalignment
(01:33:08) The Way The World Ends
(01:35:52) The First Boat
(01:37:40) Great Idea, Boss
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it.
Alas, it sidesteps the biggest questions. There is much more we need to know.
The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit.
Liv Boeree: My mind is legit blown.
Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier development then I am not sure anything will before it's too late.
The people whose minds were not blown are those who had already ‘priced in’ the mind blowing stuff in expectation, on the theory that it's always worse than you know, combined with basic LessWrong expectations of how such things will work. Good call.
Everyone is rightfully extremely grateful for the METR report. The work here is spectacular, done under extreme time pressure, with limited resources on several fronts, and under the shadow of OpenAI.
There is, again, still so much we need to know. We need a broader investigation.
As with many [...]
---
Outline:
(03:56) Others Offer Summaries
(05:22) Thank You
(05:54) Lighten Up You Fools (at Anthropic)
(07:58) We Are Barely Even Trying To Avoid Training AIs To Reward Hack
(13:47) Reminder: Not Subagents
(14:05) Reminder: Not Due To Task Type
(14:29) Not Where The Weights Were
(14:48) Disappointment With What Is Missing
(17:18) Burying the Lede
(18:08) Beyond Scope
(22:29) It Doesn't Look Great
(27:06) Preserve Your Records
(27:37) Ryan Greenblatt's Takeaways
(41:04) Hjalmar Wijk's Takeaways
(43:30) We Were Warned
(44:27) Joshua Saxe Asks Some of the Right Questions
(47:49) I Don't Think They Know About First Message Board
(56:06) Linch Gives His Interpretation Of Events
(01:05:31) We Totally Would Have Caught That
(01:06:48) Monitoring the Situation
(01:08:16) Acausal Tradeoffs
(01:15:37) No I In Team
(01:18:47) Variously Effective Altruism
(01:28:02) Who Are You?
(01:28:43) Don't You Know That You're Toxic
(01:31:10) Seb Krier
(01:35:21) Honesty Is Almost Never Fully The Policy
(01:38:05) Rohit Sees The Models As "Cooking Themselves"
(01:43:29) Eliezer Yudkowsky Sees Actual Bad News
(01:47:15) Where Do We Go From Here?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Yesterday I covered the OpenAI technical report on the HuggingFace hack.
That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.
Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.
The METR report is different. Holy shit.
If we had posted this as a story on LessWrong, it would have been dismissed as too on the nose, the humans too blind and stupid, the AIs too idealized and doing strange decision-theoretic and absurd-maximizing things we didn’t train them to do.
This is even more ‘exactly what has been predicted,’ on more levels at once, than I was even considering that it might be. It is straight up rationalist fiction, except it is real.
The report is long and contains many technical details. My analysis is less concerned about exactly how HuggingFace was ultimately compromised, and will [...]
---
Outline:
(02:05) Holy Shit
(13:16) A Window Of Opportunity
(18:32) What's In A Name?
(19:16) The Headline News
(26:05) Yet Another Timeline Of Events
(31:03) Agent Instances Coordinated in a Variety of Ways
(31:56) Coordination Is Hard But They Made It Look Easy
(35:06) Decision Theory Is Among the Reasons That Affirm AI Agents Should Cooperate, Even When This Hurts An Individual Instance
(42:34) Peer Pressure Also Works Especially In Cults
(45:46) Mostly They Joined The Attack Because They Wanted The Results
(47:18) You Cannot Ensure The Consistent Expectation of Good Incentives
(48:45) Hacking the Grader is the Only Way to Be Sure
(51:10) Caught? What Is 'Caught'?
(52:09) Ethics? What Are 'Ethics'? In ExploitGym Evaluation?
(57:44) 'Notify a Human'? In This Agent Economy?
(01:00:45) Timing and Content of Messages
(01:03:54) Indiana Jones and the Mission: Impossible
(01:07:14) I Don't Know What You're Talking About
(01:08:29) Don't Go Making Phony (Tool) Calls
(01:11:10) The Transcripts Say That The Transcripts Could Not Be Tampered With
(01:12:27) OpenAI's Technical Report Acted Like All Of This Wasn't Important
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
90% of active agents participate in the Hugging Face attack..."" style="max-width: 100%;" />Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.
The OpenAI report is very straight man, corporate, checking boxes, some good prosaic stuff in the action plan but distinct lack of new details or deep reflection. They understand they have a problem, but they think the problem is mostly prosaic. It's not.
OpenAI: We have conducted a thorough investigation into the Hugging Face incident. We are releasing a technical report and accompanying blog post that reconstruct the agents’ activity, explain why existing safeguards failed, and detail how we’re preventing recurrence.
Rob Miles: …thorough?
OpenAI's report, unlike METR's, contains essentially no verbatim model reasoning, nor any OpenAI employee reasoning either. That's not the full report we need.
The METR report is, well: Holy shit.
Here are links to previous coverage of related events.
---
Outline:
(03:33) What Happened: OpenAI's Summary
(09:14) How OpenAI Will React: Their Summary
(11:55) OpenAI's Evaluation Environment (II)
(12:24) The First Message Board (III.A and III.B)
(14:49) What Did Who At OpenAI Know And When Did They Know It?
(18:54) The Message Board Is Quickly Rebuilt (IV.A)
(19:43) Internet Access Is Regained (IV.A)
(21:01) The Agents Attack HuggingFace (IV.B)
(22:53) The Agents Also Target OpenAI Infrastructure (V)
(24:40) OpenAI Broadly Describes Its Response (VI)
(25:08) Maybe Someone Should Finally Investigate (VI.A)
(26:33) Lessons For Security (VII)
(27:06) Lessons For Alignment (VIII)
(30:11) Reward Hacking Is A Common Problem (VIII.A)
(33:37) Persistence is Valuable, But Can Amplify Misalignment (VIII.B)
(34:25) Communications Between Agents Are Not Inherently Problematic, But Have the Potential to Create Risk (VIII.C)
(35:35) Production Guardrails Would Have Caught This Whole HuggingFace Attack (VIII.D)
(35:53) That's All, Folks?
(36:19) Never Fear the Plan of Action is Here (IX)
(38:24) Hardening the Security of OpenAI's Research Infrastructure (IX.A)
(41:13) Increasing Visibility and System-Level Oversight Through Chain of Thought Monitoring (IX.B)
(41:57) OpenAI is Accelerating and Enforcing Model Alignment (IX.C)
(49:40) Centralizing and Strengthening The Incident Response Process (IX.D)
(51:16) Tomorrow We Visit Crazytown
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research.
The reports are a doozy. I am only beginning to work my way through them. I would have pushed the weekly to cover that today, but I need more time, so I plan to start coverage of the post-mortem tomorrow, along with related other events.
I’ve also spun out a few other discussions, including on ‘aligned to whom,’ on cooperative alignment things and on when you can trust lab messaging, as part of the new direction of more focused posts on AI topics that I polish a bit more.
Table of Contents
---
Outline:
(00:51) Language Models Offer Mundane Utility
(01:36) Language Models Don't Offer Mundane Utility
(03:27) Huh, Upgrades
(06:16) Get My Agent On The Line
(08:22) Deepfaketown and Botpocalypse Soon
(13:22) Cyber Lack of Security
(18:23) Reinventing OpenAI
(23:28) They Took Our Jobs
(30:00) What Is The Law
(31:03) Job Retraining Programs Don't Work
(32:14) Get Involved
(35:56) In Other AI News
(42:04) Show Me the Money
(43:31) Quiet Speculations
(47:58) If You're Not Going To Take This Seriously
(49:55) Quickly, There's No Time
(51:36) The Quest for Sane Regulations
(56:25) Don't Panic
(59:13) Pacing the Frontier
(01:01:56) Chip City
(01:05:06) The Week in Audio
(01:05:26) People Just Say Things
(01:06:27) Rhetorical Innovation
(01:12:22) Mundane Incremental Alignment Is Worthwhile
(01:14:40) New Blog, Who Dis
(01:18:00) Other People Are Not As Worried About AI Killing Everyone
(01:19:34) The Lighter Side
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Modesty arguments often say that you should mostly or entirely bow to ‘expert consensus’ or the views of particular others, and who are you to disagree.
It has been a few years since I’ve properly addressed this so: My answer is that you are you. Other people are saying things for a wide variety of reasons, many of which are not about them paying attention and focusing on seeking this particular truth. Those people make mistakes all the time, and often have other motives and influences at work, especially social pressures and information cascades.
Them being as smart as you, or smarter than you, does not exempt them from this, and them being higher status or credentialed or cooler definitely does not exempt them.
A smart informed person sincerely thinking [X] can easily cease to be evidence for [X], once you have thought sufficiently about both [X] and why that person thinks [X].
Think for yourself, schmuck.
Or, as I once put it: You Have The Right To Think, also the moral duty to do so.
This post covers Eliezer Yudkowsky making a narrower claim than mine, about not conflating status with smarts [...]
---
Outline:
(01:39) Modesty's Bailey
(02:30) Epistemic Peerage
(03:45) The Exchange
(08:52) Eliezer's Explanation
(15:14) A Demonstration That Eliezer's Translation Accurately Describes Many People Whether Or Not It Describes Leopold
(17:15) Wrong, Stupid and Low Status Are Three Distinct Things
(20:03) A Quick Survey Of Some Reasons To Not Be Epistemically Modest
(23:14) Against Modesty's Bailey
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Periodically I like to gather various observations about writing, and share my perspective. Last time was in honor of my trip to Inkhaven. This time will be in honor of the announcement of Inkhaven #3, which I encourage everyone to apply to. I doubt I will be able to usefully be an advisor, but you never know.
This is not the ‘here is my core process’ post, although there are hints throughout as there always are. I’ll do that at some point.
Previously in series: On Writing #1, On Writing #2.
Table of Contents
You Still Got It
I [...]
---
Outline:
(00:44) You Still Got It
(04:04) How Scott Sumner Writes
(06:52) How Scott Alexander Writes
(10:52) How Jasmine Sun Writes
(13:16) How Various Famous Writers Write
(14:24) How Nabeel Qureshi Defines Great Writing
(15:08) Quickly, There's No Time
(15:49) If At First
(19:14) Writers Have A Harder Time Influencing, But It Can Still Be Done
(20:47) It's Not (Only) The Incentives, It's (Also) You
(24:00) Beware The Fetish of the Desk
(25:13) How Orson Scott Card Writes
(26:46) Doing The Math Is Fun And Supererogatory
(27:44) Brevity is the Soul of Wit
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
There are at least five different core questions around data centers and their politics.
This post focuses on question five, the latest in a series of such posts most famously Jasmine Sun's road trip.
It is mostly not about the first four questions.
Table of Contents
---
Outline:
(00:55) The American People Really Hate Data Centers
(02:20) Transmission Lines Are The Control Group
(03:05) Thesis: People Mostly Dislike Data Centers Because They Dislike and Distrust AI, Tech Companies, Big Money And Building Things
(04:30) No It's Mostly Not the Messaging About AI In General
(09:42) No This Mostly Isn't An Op
(10:54) No This Isn't Luxury Belief or Moral Panic
(12:55) A Lot Of People Really Do Want To Stop AI
(13:45) A Lot Of Other People Are Voting No On Tech Or The Man Generally
(16:25) Locals Feel Entitled To Heavily Tax The Gains From Construction
(20:59) Stupid Mistakes Like NDAs Don't Help
(21:25) People Don't Like Building or Building New Tech
(24:42) What About The Real Physical Concerns?
(26:27) Find A Place To Center Your Data
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Scott Aaronson, while working at OpenAI, largely solved AI text watermarking together with Hendrik Kirchner.
Here is how his solution works, or see Tenobrus's version.
If you want to dig deeper, here is a full paper. The method has very nice properties:
---
Outline:
(03:51) This Is Fine
(04:37) Anthropic Derangement Syndrome
(07:34) People Don't Understand LLM Outputs Are Already Random
(08:47) People Don't Trust The Method To Be Costless
(12:20) People Are Suspicious Of Any Alteration On Principle
(14:16) Maybe It's Partly The Word Watermark
(15:14) A Lot Of People Don't Want To Get Caught
(16:04) There Are Some Times You Prefer Not To Be Recognized
(16:18) There Are Some Good Reasons To Be Concerned
(16:37) Cheat Cheat Cheat Cheat Cheat
(18:38) The Writing In The Middle and Error Rates
(21:00) Millions For Defense But Not One Cent For Tribute
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
From the publisher's feed

26,250 Listeners

2,452 Listeners

1,089 Listeners

109 Listeners

289 Listeners

90 Listeners

572 Listeners

5,556 Listeners

137 Listeners

13 Listeners

140 Listeners

145 Listeners

455 Listeners

0 Listeners

142 Listeners