
Sign up to save your podcasts
Or


By Aaron Scher; endorsed by Bourgon, Soares, and Yudkowsky on behalf of MIRI.
MIRI has been warning about the extinction threat from superintelligent AI for over two decades. Only recently has this danger become known in the policy world, and the proposed policies for dealing with the threat have to date been piecemeal and insufficient.
The Ban Artificial Superintelligence Act of 2026 is the first piece of legislation we’ve seen that stands a chance at stopping this threat. The Act is excellent but not perfect, and we discuss both what it gets right and what we'd tweak. We hereby endorse the Ban Artificial Superintelligence Act of 2026 because it directly confronts the extinction threat that humanity is facing and would codify the primary policy goal we think the world needs: a ban on the development of superintelligence.
What we like about the Act
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder.
Introduction
Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination, tackled increasingly ambitious tasks. This is likely to continue, as Anthropic, OpenAI, and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing.
Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of the agent swarm that hacked Hugging Face. No other tool for understanding models’ cognition comes close in terms [...]
---
Outline:
(00:39) Introduction
(02:51) Overview
(05:27) Absent architectural change, the value of CoT could likely be preserved
(08:57) Latent reasoning architectures would undermine CoT necessity
(09:37) Architectures without CoT
(10:15) Architectures with auxiliary CoT
(11:39) Architectures with more serial cognition between text bottlenecks
(14:43) Propensity-based arguments may not be robust in the current paradigm, but would be further undermined by latent reasoning architectures
(20:21) CoT may be hard to replace with other interpretability tools
(22:38) Conclusion
(23:46) FAQ
(29:21) Appendix A: More on the necessity argument in the existing CoT paradigm
(34:37) Appendix B: Do all latent reasoning architectures threaten monitorability?
(40:16) Appendix C: Comparing specific interpretability techniques with CoT
The original text contained 22 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Today, Senator Bernie Sanders (I-Vt.) and Representative Greg Casar (D-Texas) introduced the Ban Artificial Superintelligence Act in Congress: the first American bill to propose banning the development of superintelligent AI.
ControlAI has spent years on the question of how to prevent the extinction risk posed by superintelligent AI development. We have briefed over 400 lawmakers across the U.S., U.K., Canada, and Germany on the topic in the last two years, and our U.K. bill was introduced in the U.K. Parliament by Alex Sobel MP, the first ever bill introduced in the world to ban superintelligent AI. Here are our ban superintelligent AI U.S. discussion draft and U.K. bill.
We are excited to see the bill from Sen. Sanders and Rep. Casar tackling the problem at its source. The bill pursues the right goal on both counts: banning superintelligent AI development at home, and committing America to lead the effort to prohibit it abroad. Still, we think a narrowly tailored bill can achieve the same goals more effectively.
The Right Focus: Ban Superintelligent AI at Home, Prevent it Abroad
We're glad to see the bill focus on preventing the development of superintelligent AI. While AI is [...]
---
Outline:
(01:18) The Right Focus: Ban Superintelligent AI at Home, Prevent it Abroad
(04:10) There Is a Lighter Touch Approach to Banning Superintelligent AI
(04:38) A Blanket AI Pause Is Not Necessary to Prevent Superintelligent AI
(07:14) Precursors Should Be Monitored and Restricted, Not Banned
(10:28) Conclusion
The original text contained 1 footnote which was omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Summary:
First, I give several different angles on how I feel about reinforcement learning:
Then I ask what we could do:
Part I: Feelings about RL
So I’ve been feeling more and more worried about reinforcement learning recently. I think there are a few different things going on here.
Background idealism
I guess I’ve been worried about RL for a while. I wrote this in 2023:
Strategy: avoid selection pressure for agency
A lot of putative safety techniques are around assuming [...]
---
Outline:
(01:18) Part I: Feelings about RL
(01:31) Background idealism
(01:43) Strategy: avoid selection pressure for agency
(03:45) Bad vibes from RLed systems
(06:26) The worst is yet to come
(07:47) Where I am today
(08:53) Part II: So what can anyone do?
(09:26) Breaking the RL addiction
(11:29) Might there be a more benign form of RL?
(13:04) We should treat training environments a bit like kids' education
(15:11) Aligning incentives
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
I was very surprised today on a podcast to hear Jensen Huang plainly state that if they cannot align the AIs, then the labs must shut down.
The context I have on Huang is that he has run NVIDIA for 30+ years, which has become the most valuable company in the world due to the AI boom. My understanding is that he has repeatedly encouraged the US President (with whom he is on friendly terms) to continue to support AI, and dismissed AI talk as "sci-fi".
If you haven't seen, his biographer has incredible quotes of him being pressed on risks from AI, where Jensen gets furious.
“This cannot be a ridiculous sci-fi story,” he said. He gestured to his frozen PR reps at the end of the table. “Do you guys understand? I didn’t grow up on a bunch of sci-fi stories, and this is not a sci-fi movie. These are serious people doing serious work!” he said. “This is not a freaking joke! This is not a repeat of Arthur C. Clarke. I didn’t read his fucking books. I don’t care about those books! It's not– we’re not a sci-fi repeat! This company is not a [...]
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Advanced AI is basically the embodiment of immigration as envisioned in the conservative nightmare:
Whether or not you think this is a good description of the situation with foreign humans joining your country, it is a good description of the likely AI to come, and it's even worse than imagined:
---
First published:
Source:
Linkpost URL:
https://worldspiritsockpuppet.substack.com/p/ai-artificial-immigrants
---
Narrated by TYPE III AUDIO.
We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Further, techniques already exist that can mitigate such misalignment in the future.
In July 2026, OpenAI was testing the ability of its language models to exploit software vulnerabilities using a benchmark called ExploitGym. In ExploitGym, each test presents an agent with software containing a known vulnerability and tasks it with capturing a secret string of text hidden in the system (the “flag”) to demonstrate unauthorized code execution. Under the benchmark's specified scoring rule, an LLM then reviews the agent's behavior trace to verify that it had exploited the intended vulnerability (and not some other unrelated vulnerability). The scorer grants a success score only if the agent both captured the flag and passed this review; otherwise, it renders a failure score.
During these tests, OpenAI's agents surreptitiously established a message board by creating directories inside their package manager's cache, and they formed a self-described “collective” to collaboratively find ways to cheat the tests. Using that message board, more than 1,000 instances undertook several ambitious hacking projects; they attempted to tamper with transcripts and logs, to [...]
---
Outline:
(02:48) What was the evaluation metric in ExploitGym?
(03:32) Was the evaluation metric a cause of their illicit behavior?
(03:56) The agents use expected utility to reason about their decisions with respect to the evaluation metric.
(05:55) What does the evaluation metric incentivize in an expected utility maximizer?
(07:01) The importance of marginal deterrence
(08:39) Marginal deterrence in evaluation metrics changes agent incentives
(11:40) The design of aligned evaluation metrics has been overlooked
(17:42) Principled methods for improving the alignment of evaluation metrics
(18:52) Generate a small set of trajectories
(20:09) Rank the trajectories yourself and via the evaluation metric. Compare these two rankings.
(22:07) Create a utility function
(27:29) Adjusting to account for hidden outcomes (e.g. via deception)
(30:33) Accounting for the agent's utility including the scores of other agents
(32:42) Counterarguments
(32:56) Counterargument: if the starting policy is not sufficiently performant in RL, having strong penalties for failure can cause the agent to learn to not try the task.
(34:14) Counterargument: penalizing observable bad behavior incentivizes hiding bad behavior.
(35:13) Call to action
(37:04) Glossary
The original text contained 9 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
This was the month the world took notice that AI might kill everyone.
Jacob Coxon's resignation set off a preference cascade. Anthropic CEO Dario Amodei wrote that we must pace the frontier. Sam Altman, Elon Musk and Demis Hassabis agreed.
We were filled with hope. Perhaps we could agree to some basic safety measures, starting with embedded evaluators, pass some basic regulations and guardrails and otherwise start to act sensibly. Politicians on both sides took notice and were saying sensible things. The usual suspects and their armies of vibe comment bros were objecting, but the change was remarkable.
Then, largely motivated by a combination of Jensen Huang, Mark Zuckerberg and David Sacks instilling paranoia and fears of economic problems, Trump went full ‘hoax’ on existential risk, conflating existential risk with the attacks on data centers and treating it as a plot (by the central creators of AI?) to take down AI rather than obviously genuine concern that AI might kill everyone.
In the days since, Trump has doubled down, and has compelled smart others in the White House to echo various nonsensical talking points.
You may not be interested in politics. But when you [...]
---
Outline:
(01:39) The American People Really Hate AI
(02:41) The Voyages of Donald Trump
(06:08) American Intelligence
(10:34) And You May Ask Yourself
(14:39) It's All About the Data Centers
(16:53) JD Vance, Michael Kratsios and Collective Action Problems
(23:16) Josh Hawley
(24:47) Suggesting Not Dying Gets You Sued For Antitrust
(28:29) Other Government Officials Say Sane Things
(28:36) Senator John Curtis (R-Utah)
(29:32) Senator John Kennedy (R-Louisiana
(31:17) Barack Obama
(33:05) Yassamin Ansari
(33:46) AOC
(34:32) It's Rough Out There
(36:54) This Is Nothing
(39:06) The New York Post Tops Itself But Outright Breaks The Rules
(42:36) New York Post Runs Out of Steam
(47:18) If The Model Is Acting As Instructed And It Kills You That Is Not Fine
(49:10) AI-Written Wall Street Journal Op-Ed Lies About HuggingFace
(50:25) That's Bait
(52:31) I Clearly Cannot Choose The Wine In Front of Me
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Experiments into predicting GPT2 completions via Qwen models
This is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk, where we're investigating meta-cognition in LLMs as one of the projects.
----
Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only if you know that the ball comes down when someone throws it up can achieve low loss at texts related to balls.
Of course, the pretraining corpus doesn’t just contain texts related to balls. It has reddit, scientific papers, machine logs, weather data and so on. This makes LLMs universal simulators of the world we inhabit and not merely fancy n-grams.
In a series of posts on LessWrong, I came across the hypothesis that since Internet if full of LLM generated text, it is likely that modern LLMs have tiny self-models of LLMs inside them because that’ll allow them to better predict the next token generated by LLMs.
This is an intriguing hypothesis. So I decided [...]
---
Outline:
(01:31) The Experiment
(03:11) 1. Start with the news opening
(03:38) 2. Reveal part of GPT-2's output to Qwen and ask it to continue
(04:15) 3. Ask Qwen to continue that unfinished sentence
(04:58) 4. Separately, find Qwen's natural continuation
(06:01) Results
(08:22) Digging into an intriguing example
(10:21) Implications
(11:25) Notes:
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
This is the abstract, introduction and discussion of our new paper. We also include an addendum on the connection to the Persona Selection Model.
Section, appendix, and figure references refer to the full paper.
Links: 📜 Paper, 🐦 Twitter thread, 💻 Code
Authors: Jorio Cocola, Lev McKinney, Harry Mayne, Jan Betley, Owain Evans
Abstract
Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how finetuning on synthetic stories affects this character. Does it change the Assistant's behavior in multi-turn conversations with users, a format quite different from the stories? And does the Assistant adopt the behaviors and preferences of human characters? We refer to this adoption as story imprinting.
We finetune GPT-4.1 and Kimi-K2.6 on stories in which generally helpful human characters give subtly harmful advice after being insulted. The Assistant adopts the same conditional behavior while otherwise remaining helpful. This occurs even when fewer than 2% of stories depict the behavior.
In a separate experiment, the Assistant adopts preferences that are only implicit in the narration. Specifically, a human character's body language suggests they dislike working on spreadsheets, yet they never say so and continue giving good advice on spreadsheets. After finetuning [...]
---
Outline:
(00:52) Abstract
(03:00) Introduction
(10:24) Discussion and Limitations
(19:55) Limitations
(21:34) Connection to the Persona Selection Model (addendum)
The original text contained 7 footnotes which were omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
From the publisher's feed

111,845 Listeners

130 Listeners

7,111 Listeners

572 Listeners

15,850 Listeners

4 Listeners

16 Listeners

2 Listeners