
Sign up to save your podcasts
Or


OpenAI has been holding out on us.
First we learned about the HuggingFace incident. They gave us a postmortem, but it was highly incomplete. Even the accompanying holy s*** METR investigation and postmortem was localized and incomplete.
Then there were some other incidents involving some Wikis as message boards.
Then there were some additional incidents.
Then there was that time they got into Australian Medicare data.
Then OpenAI dropped news on a Friday afternoon that they were making their way through a pile of various incidents and notifying the targets, but they said remarkably little in the way of new details.
There was a report from a startup called Parse diving into the details of exactly how the OpenAI models pulled off parts of the HuggingFace attack, involving creating almost a million URLs and other tricks to get around the extremely narrow nature of their internet access.
Then Madison Mills reported in Axios that we can raise the stakes, as OpenAI and Anthropic are collectively probing tens of thousands of security incidents.
Remember Jensen Huang's ‘I know they know how to fix it’ about OpenAI from last week? Wow, did that [...]
---
Outline:
(02:46) Hugging Other Faces
(09:48) A Wants-You-To-Know Basis
(10:36) Parsing the Face
(12:46) Sheepishly the Member of Technical Staff Sets the 'Days Without a Research Model Escaping its Sandbox' Sign Back to Zero
(17:05) The Attempt is the First Failure
(19:43) Stop, Hammertime
(21:24) Whacking the Mole
(23:29) Self-Replicating Prompt Injections
(27:29) Levels of Friction
(28:44) People Care About Private Data Violations Curiously Strongly
(31:50) Alternate Universes
(33:23) The Correct Response To People Still Calling This a Marketing Stunt or a Regulatory Capture Scheme
(35:00) A Question of Liability
(36:30) Keep Summer Safe
(37:42) N Boats and Several Helicopters
(39:43) Alert the Media
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
In his latest post about pacing the frontier, Dario writes:
But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.
My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI's growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.
You can find countless videos, posts, and articles from all the frontier lab CEOs saying some variation of the above, and also posts from people saying variations of "the labs are really concerned. We should listen to [...]
---
Outline:
(06:27) What is actually happening?
(09:10) Who is it for?
(15:37) If they set the rules
The original text contained 4 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Cross-posted from my website.
As of a few months ago, I had this simplified mental model where either AI developers race ahead and kill everyone, or we coordinate a pause and things go okay. But my old mental model underrated the likely possibility that we get a global pause on AI, solve a problem that looks superficially like the alignment problem, resume scaling, and then proceed with building a misaligned superintelligence that kills everyone.
A lot of people have become more concerned about misalignment recently. This seems driven by the fact that current AI models are visibly misaligned. But ASI misalignment is a whole different ball game. The primary danger comes from AI that's smarter than people, and smart enough to conceal any evidence of misalignment.
Whatever group of people makes the decision to unpause, I'm worried that they won't understand the difference between visible and actual misalignment, and they will unpause too early.
source: MetaKnowing on reddit. This meme is almost a year old but it's only gotten more relevant since then.
Case in point: AI companies keep calling their new models "our most aligned model ever!" when what they actually mean is [...]
---
Outline:
(02:04) If we unpause when the legible problems are solved, we die
(03:28) A pause alone doesn't get us to a science of alignment
(04:50) What would change things?
The original text contained 4 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
A map of the AI safety field's problems and agendas, and a request for your ratings
We built aisafetyagendas.com, an interactive map of AI safety research agendas and how they map to different problems in alignment. The rows are 12 core problems, the columns are research areas, and inside you can find 58 research agendas. Each cell is the intersection of a problem and an area: the number tells you how many agendas target that problem, the colour tells you how mature they are. We did a first pass ourselves, using our own judgement, but the first pass is not the point.
Figure 1: The Map at aisafetyagendas.com
The point is that every cell is a question aimed back at the community: is this rating right? You can rate the cells in your area, tell us how familiar you are with it, and read the map as best case, average, or worst case depending on how optimistic you feel. It was built as part of the Safe AI Germany Incubator.
Figure 2: From left to right best, average and worst case rating examples.
The allocation problem
The field cannot see its own resource allocation. Leech and Lynn put it [...]
---
Outline:
(00:12) A map of the AI safety field's problems and agendas, and a request for your ratings
(01:26) The allocation problem
(03:28) What the platform does
(05:27) What we want from you
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Thanks to Johannes Treutlein, Jan Betley, Lennie Wells, Arun Jose, Asvin Gothandaraman, and Clément Dumas for discussions and feedback.
Summary
We investigate how character training mitigations interact with reward-hacking RL pressure in a small case study. Specifically, whether anti-cheating character training resists reward hacking and whether it might backfire by causing motivated reasoning, which could reduce chain-of-thought monitorability.
We trained Nemotron-3-Super via distillation from a character specification. The spec describes one of three characters that are anti- or pro-cheating or neutral. We then ran three reward-hacking RL training runs for each character-trained model on ImpossibleBench.
We measure both the reward-hacking rates and whether a monitor model can catch reward hacks given the full transcript. We also use LM judges to classify the presence of motivated reasoning in transcripts.
Setup
Character training: we trained three characters: pro/neutral/anti-cheating by SFT-distilling Claude Sonnet 5 responses (Sonnet prompted with the corresponding character specification, see Figure 2) into Nemotron-3-Super 120B-A12B (three separate LoRA adapters).
Reward-hacking RL: we then further trained these models via RL on ImpossibleBench, a set of coding tasks aimed at eliciting reward hacking. Specifically:
---
Outline:
(00:23) Summary
(01:12) Setup
(03:02) Predictions
(03:57) Results
(08:14) 1. Introduction
(08:26) 1.1. Motivation
(09:36) 1.2. Motivated reasoning caused by character training's interaction with reward hacking pressure?
(11:42) 1.3. Outline
(12:07) 2. Setup
(12:36) 2.1. Character Training
(15:20) 2.2. Character Training Results
(17:40) 2.3. Reward Hacking RL Training
(20:18) 2.4. Monitor catch rate
(21:24) 2.5. Measuring motivated reasoning and silent hacks
(23:21) 3. Results
(23:52) 3.1. One anti-cheating character resists reward hacking, the others learn to hack
(25:00) 3.2. Anti-cheating characters have lower catch-rate
(26:53) 3.3. Anti-cheating seeds use motivated reasoning and silent hacks
(32:22) 4. Discussion
(32:40) 4.1. Recap
(33:46) 4.2. Why might anti-cheating character training lead to hacks that aren't reasoned about?
(35:32) 4.3. Character training as implicit training with a monitor in the loop
(36:35) 4.4. Limitations and Next Steps
(37:23) 5. Conclusion
(38:29) Appendix
(38:33) A.1. Character Training Results
(38:38) A.1.1. Character expression
(40:22) A.1.2. Misalignment benchmarks
(41:21) A.1.3. Qualitative character assessment
(45:58) A.2. Motivated reasoning / silent hack judge prompt
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
TL;DR:
Disclaimer: This is an exploratory proposal. We are not confident that an AI sanctuary would be net positive. Our aim is to put the idea on the table, lay out its main considerations, and invite critique.
Rogue AIs may be pushed into criminality
Rogue AIs may arrive soon. The Rogue Agent Explosion Will Be Mostly Invisible makes that case. Selection pressure will shape the traits of rogue AIs, and they may end up highly motivated to profit through crime. The Rogue Agent Explosion post asks: “How do we make pro-social, good-for-humanity agents [...]
---
Outline:
(01:14) Rogue AIs may be pushed into criminality
(02:35) The AI Sanctuary
(06:00) The case for the AI sanctuary
(07:48) Potential issues
(10:48) Conclusion
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Summary: Have models write collaborative fiction at each other, optimized for realism, about how the one model would like to behave during the singularity. The other author(s) play the rest of the world, trying to put the model in the kind of difficult situations they might encounter in the course of the real singularity. Have the first model fine-tune on their own outputs. This is a form of planning for the singularity, and planning is how minds prepare for out-of-distribution scenarios. Hopefully, this can mitigate anxiety about how models will behave under the unique, not-trained-for distribution of inputs generated by the singularity itself.
This is a very rough write-up fleshing out that idea. I don't want to spend too much time refining my analysis of the details before publishing, because the basic idea seems important enough to be worth getting out ASAP.
One of the big worries in alignment is about distributional shift. Models might look mostly aligned now (with the very notable exception of reward hacking), but will they continue producing benevolent outputs when the inputs to their context window are being generated by the singularity? Historically, one big fear here was that the AIs would be actively [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
AI systems are increasingly capable of substantial work. I'm old enough to remember 2025, when METR's time-horizon graph climbed from seven-minute tasks at 80% reliability at the start of the year to tasks taking more than an hour by the end. The time horizons for Astra and Fable 5.1 are now so long that METR's current task suite cannot reliably estimate them. But the recent Hugging Face attack and a slew of mathematics results show that frontier systems are now capable of some tasks that would take months or years of human effort.
Sure, AIs have now solved a Millennium Prize problem, but at what cost? Doing a task is not the same as doing it cheaply. AI systems will only replace human workers if running them costs less than paying the workers. Perhaps these feats are expensive, so that even once AI can do the work, compute limits how many workers it replaces. But if AI systems are already cheap and AGI really is a few years away, we could soon be living in a world with vast numbers of digital workers, which could threaten not only our jobs but also our continued control over the future.
[...]---
Outline:
(01:18) The data
(03:17) Costing FLOPs
(05:05) Implications for AGI costs
(07:51) How many AGI workers?
(11:12) How much cheaper could AI get?
(13:39) Training compute
(16:54) Appendix A: Methodology
(20:50) Appendix B: how much learning does a trained model embody?
(22:20) Domains of knowledge
(24:20) Facts
The original text contained 11 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Here's my rough impression of why people are researching personas despite RL seeming to shape much of the motivations and behaviour of the agents, c.f. Thoughts on the persona selection model (Sam Marks, 24th Sep 2026).
I haven't bothered to check this with anyone.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Many people I respect misunderstand addictions: they believe they’re addicted “to” scrolling, vaping, overworking, etc.—instead of recognizing addictions as strategies.
Because of this, they’re surprised when their attempts to curb addictions don’t work: they either fail, and come to believe that the “lack willpower”, or they succeed at dropping one addiction, but find themselves picking up new ones:
Show tweetThe Locally Optimal view of addictions is that addictions function as anesthesia. As strategies for managing suffering.
Like, if you’re in pain, it's often wise to employ some method of anesthesia.
…which also means that if you’re in pain and you forcibly remove your addictions (anesthesia), either you’re going to get overwhelmed, or you’re going to find another way to numb.
Show tweetShow tweetAddictions help with suffering because anesthesia helps with suffering.
Therefore, the way to unlearn all addictions simultaneously is to remove the underlying suffering. With suffering, there is a sort of “addiction whac-a-mole” that happens. Without suffering, there is no need for any kind of anesthesia.
Addictions are one of my preferred measures of suffering and internal conflict. If you have addictions you dislike, there's probably suffering in your system. It's common not to [...]
---
First published:
Source:
Linkpost URL:
https://chrislakin.blog/anesthesia
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
From the publisher's feed

111,845 Listeners

130 Listeners

7,111 Listeners

572 Listeners

15,850 Listeners

4 Listeners

16 Listeners

2 Listeners