
Sign up to save your podcasts
Or


We are in the midst of a preference cascade about existential risk from AI.
A preference cascade is, alas, the best method we have to change the debate.
The avalanche has started. There is still time for the pebbles to vote. For now.
Mike Solana gave the correct view of why Coxon's post went viral, which is that enough Americans finally have enough context on AI to care, and there were enough big accounts that were happy to amplify the Tweet quickly to get it initial attention. That is all you need when there is enough dry tinder.
What we must realize is that the current preference cascade, on the need to Pace the Frontier, is insufficient. If we are to make it out of this alive, we will have to do better. We have to, as Dan Selsam warns, actually solve the underlying problems.
The next step is to continue the cascade. That includes inside the labs, and also among the media and politics. It includes both people who previously focused on other things stepping up and new voices being heard.
A lot of that will be overcoming the inevitable political opposition [...]
---
Outline:
(01:38) The Cascade Was a Long Time Coming
(02:59) The Cascade Has Reached The People
(04:38) Elon Musk Doubles Down
(05:18) Matthew Yglesias Steps Up
(08:42) Op Eds and Posts Are Written
(11:05) Jacob Coxon AMA
(17:35) Bilal Chughtai Quits DeepMind and Sounds the Alarm
(20:29) The Cascade Is Insufficient
(21:47) What Would It Take
(28:10) OpenAI's Dan Selsam Sounds A Louder Alarm
(41:59) Some People Worry On Meta Levels You Never Imagined
(43:10) Two Kinds of Threats
(44:41) The Two Towers and The Narrow Path
(46:45) A Specific, Detailed Story About AI Killing Everyone That Doesn't Sound To Me Like Science Fiction
(50:06) What Can I Do About It?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
You have a horse.
You do not like the horse. The horse does not like you.
At the moment, you are completely dependent on the horse. The terrain is impossible to traverse on foot. There is no way to travel without a horse. You wish that would change, but when you tell other people, they laugh and call it impossible. A few get angry.
You must spend hours each day feeding, cleaning, and taking care of the horse. You must spend even more time working to earn enough money to pay for the horse's needs. The horse is often unsatisfied with your offerings. No matter how expensive and time consuming your efforts, the horse will desire something more unique, exciting, or comforting.
The horse also requires a third of your day to sit and do nothing. During this time, you cannot do anything or go anywhere.
If you do not comply with the horse's desires, it will make your life miserable.
People tell you that, as the horse's rider, you have complete control over the horse. Somebody must have forgotten to tell the horse this. If the horse is hungry or thirsty, it will draw your [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
While this is relevant to my work at MIRI, I have not checked these ideas with anyone else on the team and am posting this on my personal LW account. These views are my own. And to be honest, I am writing this mostly to remind myself of my weakness.
---
I expect one (or many) adversarial memetic attacks aiming to trip you up, perhaps consisting of fake leaks relating to dangerous stuff happening in the labs. Specifically, worrying incidents that may fit snugly within your worldview, leaking from multiple sources including news outlet/s, but not confirmed/confirmable by a primary source. Think rumors about exfiltrated weights, AIs attempting to create viruses, agent swarms hacking into and gathering information from nuclear infrastructure, etc.
An easy way to remove status from a movement is to trip it up: make it fall for a misinformation trap in public, then use that slip-up to discredit the movement for all time. The game is set for a memetic attack like this. There's a well-resourced group waiting for your screw-up.
And then you may remember much that will help you.
In public and in private, if you feel surprised or confused, notice your confusion. These [...]
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Three whitehack hackers from Hacktron used Claude Opus 5 within hours of release to chain exploits into hacking to OpenAI's monorepo codebase.
This likely means they have access to almost all of OpenAI's research and production code, though likely not the literal model weights. Oops.
You can so their blog post about it here. Interesting sidenote: they used less than $3000 of compute credits for the entire hack.
Alternative title: OpenAI unilaterally implements "Total Research Transparency" from Plan A.
---
First published:
Source:
Linkpost URL:
https://www.wsj.com/tech/ai/hackers-used-anthropics-claude-to-break-into-openai-b40ba883
---
Narrated by TYPE III AUDIO.
We use RL to teach a deep recurrent model and a normal CoT model to solve a math problem while hiding from a CoT monitor which of two possible problems it is solving. The deep recurrent model moves its reasoning into latents (thereby evading the CoT monitor) within 40 steps of RL; the CoT control is much worse at learning to confuse the LLM monitor. Note: in the above graph, “monitorability” is how much better than chance the CoT monitor is at figuring out which problem the model is solving. This was our simplest setting; similar settings show similar results with somewhat smaller differences between the CoT model and the deep recurrent model, see Results.
Some proposed architectural changes add a latent channel in addition to the normal text channel for persisting state between forward passes. These parallel latents architectures (like T2MLR and full-bandwidth transformers) could theoretically perform all their reasoning in latents, rather than text. We’ve speculated that these architectures will therefore be harder to oversee. This post empirically investigates whether these models can learn to reason without their CoT (which would likely make them harder to oversee) in a simplified setting.
We compare the performance of a [...]
---
Outline:
(03:49) A formal description of the parallel latents architecture
(05:00) Experiment 1
(05:50) Tasks
(09:09) Stage 1: SFT
(11:53) Stage 2: RL
(13:06) Results
(20:50) Discussion
(21:55) Task
(22:50) Pressure to solve the task
(23:32) Oversight method
(24:25) Model
(25:53) Obfuscation pressure
(28:04) Suggestions for future work
(29:58) Conclusion
(30:26) Acknowledgements
(30:49) Appendix
(30:52) Experiment 2: Output Spillover
(32:22) Result
(33:25) Experiment 3: Training pressure variations
(37:09) en-US-AvaMultilingualNeural__ When starting with low task performance, rejection filtering on highly monitorable trajectories can lead to increased obfuscation.
(37:20) en-US-AvaMultilingualNeural__ When initial performance on the task is high, rejection filtering does not exert significant pressure on monitorability.
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
I now consider it plausible that some form of recursive self-improvement is imminent, and that we may be on track for superintelligence by Christmas of this year if racing continues.
This is substantially faster than any forecast, including ones like AI 2027 that were considered outrageously fast a year ago. It is faster than I myself expected even a week ago. I don't work at a scaling lab. I don't know more than is public knowledge.
Let me be perfectly clear: what I am saying is absolutely nuts. Extraordinary claims require extraordinary evidence. I claim we have now received said evidence and you should update accordingly.
FOOM should probably should be your *default expectation*.
People have strong status quo bias. Your default expectation should be that things will radically speed up.
We are not at the ceiling of intelligence.
We should probably expect the transition to superintelligence to be incredibly fast.
RSI is a positive feedback loop, so it is inherently (hyper)exponential. Everything is an S-curve eventually, but nothing suggests the ceiling is anywhere near human level, or that it happens at a human timescale.
AI is [...]
---
Outline:
(01:08) FOOM should probably should be your *default expectation*.
(01:44) AI is capable of revolutionary advances in mathematics. Machine learning research is not different in kind.
(03:16) The speed of AI progress continues to be underestimated; by superforecasters and even by the researchers themselves.
(05:07) Internal models are significantly ahead of released ones;
(06:08) Intuitions about timing from pre-training runs are misleading since most progress comes from RL, unhobbling and algorithmic innovations
(06:29) Enter the Swarm
(07:04) Anthropic's own report states it has 30,000 agents running concurrently, and Claude has completely taken over 26% of all R&D.
The original text contained 3 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
The AI Risk grantmakers do not act like they believe in imminent existential risk from AI
The idea of "revealed preferences" is one of the most useful in economics; it allows us to cut through a great deal of metaphysical angst about what someone "really" believes, and focus on what they act like they believe, which is much more useful for making predictions about their future actions. As one example, I grew up in a, shall we say, fervently-religious community, and it's often hard for nerdy Rationalist types to understand this, but: there are people who genuinely believe in Hell, and in Heaven. They genuinely believe that moving souls from one to the other is the most important thing on Earth. It's one thing to doubt the conviction of someone who lives an easy, staid, middle-class life... but for others, their choices and behaviors (e.g. years-long missionary trips) reveal their true preference and/or belief beyond any reasonable doubt.
I bring this up because, per the actions and decisions of grantmakers operating in the AI Risk space, they mostly DO NOT seem to believe in imminent existential risk of AI. On the contrary, they act like people who [...]
---
Outline:
(00:10) The AI Risk grantmakers do not act like they believe in imminent existential risk from AI
(01:27) The explore-exploit tradeoff
(02:44) The evidence we're in "exploit" mode
(02:49) Exhibit A
(03:09) Exhibit B
(03:32) Exhibit C
(04:32) Obvious verdict is obvious
(06:35) Explore mode: Just do (good) things (better)
(09:35) Conclusion
The original text contained 13 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Below is the executive summary from our new paper at pacing.tech. The full paper is available on the site and as a PDF. The full author list is Raymond Douglas, Charles Dillon, Nikola Moore, Gavin Leech, Shahar Avin, Mathias Kirk Bonde, Rohit Krishnan, Noah Perez, Nathan Young, Cormac Slade Byrd, Stephen Casper, Jan Kulveit, & David Duvenaud
“Pacing AI” usually refers to how to conclusively handle the most extreme risks in the face of race dynamics. However, even for the goal of handling these highest-stakes cases, it's useful to take a broad view of pacing—one that encompasses all interventions aimed at moderating the pace of AI development, deployment, or diffusion. Thus:
---
First published:
Source:
Linkpost URL:
pacing.tech
---
Narrated by TYPE III AUDIO.
Summary
I was a fellow at plzdontkillus, a month-long creator bootcamp at Lighthaven, partially funded by MIRI, where ~55 fellows posted one video per day.
plzdontkillus.com originally claimed “21M+ AI risk views” with no breakdown. After I shared a draft of this post, the organizers relabeled it “X-Risk Relevant Views” and published one.
Three videos account for 80% of the views: a datacenter-water-use debunk (8.5M), an AI dystopia video (6.4M), and a Rob Miles Hugging Face incident explainer (2.5M). The rest total 4.3M. Under my stricter definition of AI safety content, fellows generated ~2 million views total.
Based on my analysis, fellow-made AI safety videos made up around ¼ of fellows’ output and ~2% of total views. 13 out of ~55 fellows posted zero AI safety videos, and an additional 8 posted only one or two.
This is partly because the program didn't incentivize AI safety content. If they run it again, I think they should change that.
Me
I’m Josh Thor.
---
Outline:
(00:13) Summary
(01:28) Me
(02:12) What they claim
(05:26) My analysis
(06:58) Program incentives
(08:38) Aella's response
(11:00) My recommendation
The original text contained 11 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
TL;DR. In this work we study obstacles to the faithful automation of alignment research. We see this as a scalable oversight problem. There are plenty of examples of how models fail at this, and as models become more capable our ability to notice these failures will diminish: even the best human checkers won’t be able to tell if the model was well elicited, thorough checking will become too costly, and models could tailor their responses to their judges. We draw on empirical examples from Geoguessr and auto-alignment runs from Arcadia's internal research to make general claims about obstacles to the oversight of fuzzy alignment-related tasks. Narrowing our attention to one prominent scalable oversight method, we find that whilst debate shows promise on typical capabilities benchmarks (aligning with recent work) it fails on tasks involving judgment calls akin to those arising in automated alignment research.
We’d like to thank David Africa, Andrew Draganov, Rory Greig, Joshua Jacob, Rishub Jain, Zac Kenton, Francis Rhys Ward and Lennie Wells for helpful feedback on this post.
Introduction
Existing empirical work on debate [1, 2, 3, 4, 5, 6] has almost exclusively focused on objective, verifiable domains, seeking to mitigate misalignment caused by supervision [...]
---
Outline:
(01:20) Introduction
(04:00) Decomposition of explanations
(10:00) Empirical Examples
(10:18) Geoguessr Setting
(11:59) Example claims in fuzzy arguments
(12:29) Nature of arguments in non-fuzzy tasks
(14:39) Discussion: scalable oversight of fuzzy tasks
(17:05) Empirical Debate Results
(17:09) Geoguessr
(18:39) LMCA Debate
(19:43) Conclusion
(20:18) Appendix
(20:21) Geoguessr Setting
The original text contained 7 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
From the publisher's feed

111,845 Listeners

130 Listeners

7,111 Listeners

572 Listeners

15,850 Listeners

4 Listeners

16 Listeners

2 Listeners