LessWrong (30+ Karma)

LessWrong (30+ Karma)

Download on the App Store

LessWrong (30+ Karma) episodes

  • “MIRI’s Position on the Ban Artificial Superintelligence Act of 2026” by Aaron_Scher

    By Aaron Scher; endorsed by Bourgon, Soares, and Yudkowsky on behalf of MIRI.

    MIRI has been warning about the extinction threat from superintelligent AI for over two decades. Only recently has this danger become known in the policy world, and the proposed policies for dealing with the threat have to date been piecemeal and insufficient.

    The Ban Artificial Superintelligence Act of 2026 is the first piece of legislation we’ve seen that stands a chance at stopping this threat. The Act is excellent but not perfect, and we discuss both what it gets right and what we'd tweak. We hereby endorse the Ban Artificial Superintelligence Act of 2026 because it directly confronts the extinction threat that humanity is facing and would codify the primary policy goal we think the world needs: a ban on the development of superintelligence.

    What we like about the Act

    • Banning artificial superintelligence (ASI), or variants of such a plan, is the only effective solution to avoid the ASI threat, at least in the near term. Most other legislative proposals do not confront this threat head-on and thus would not be effective, even if implemented. For more on why we believe this, see [...]

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/jszKCKwvzfmsNetNZ/miri-s-position-on-the-ban-artificial-superintelligence-act

    ---

    Narrated by TYPE III AUDIO.

    7 min
  • “Latent reasoning architectures would undermine CoT, our strongest oversight tool” by Lukas Finnveden, Alexa Pan, Alek Westover, Girish Gupta, frisby, ryan_greenblatt

    Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder.

    Introduction

    Swarms of more than a thousand AI agents have in recent months, both intentionally and in unsanctioned, rogue coordination, tackled increasingly ambitious tasks. This is likely to continue, as Anthropic, OpenAI, and other AI companies deploy increasingly large quantities of superhumanly fast agents to automate AI development. As the AIs increase in both number and capability, humans will find it increasingly difficult to understand what they are doing.

    Today, the overwhelming majority of our (limited) information about AI systems’ internal workings comes from (i) their CoT, and (ii) natural language communication directly between them. For example, it was only by reading CoTs and communication between agents that investigators were able to gain some understanding of the activities and motivations of the agent swarm that hacked Hugging Face. No other tool for understanding models’ cognition comes close in terms [...]

    ---

    Outline:

    (00:39) Introduction

    (02:51) Overview

    (05:27) Absent architectural change, the value of CoT could likely be preserved

    (08:57) Latent reasoning architectures would undermine CoT necessity

    (09:37) Architectures without CoT

    (10:15) Architectures with auxiliary CoT

    (11:39) Architectures with more serial cognition between text bottlenecks

    (14:43) Propensity-based arguments may not be robust in the current paradigm, but would be further undermined by latent reasoning architectures

    (20:21) CoT may be hard to replace with other interpretability tools

    (22:38) Conclusion

    (23:46) FAQ

    (29:21) Appendix A: More on the necessity argument in the existing CoT paradigm

    (34:37) Appendix B: Do all latent reasoning architectures threaten monitorability?

    (40:16) Appendix C: Comparing specific interpretability techniques with CoT

    The original text contained 22 footnotes which were omitted from this narration.

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/6m29SfjbittooYojj/latent-reasoning-architectures-would-undermine-cot-our

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    45 min
  • “The First American Bill to Ban Superintelligent AI Is Here” by Andrea_Miotti, Connor Leahy

    Today, Senator Bernie Sanders (I-Vt.) and Representative Greg Casar (D-Texas) introduced the Ban Artificial Superintelligence Act in Congress: the first American bill to propose banning the development of superintelligent AI.

    ControlAI has spent years on the question of how to prevent the extinction risk posed by superintelligent AI development. We have briefed over 400 lawmakers across the U.S., U.K., Canada, and Germany on the topic in the last two years, and our U.K. bill was introduced in the U.K. Parliament by Alex Sobel MP, the first ever bill introduced in the world to ban superintelligent AI. Here are our ban superintelligent AI U.S. discussion draft and U.K. bill.

    We are excited to see the bill from Sen. Sanders and Rep. Casar tackling the problem at its source. The bill pursues the right goal on both counts: banning superintelligent AI development at home, and committing America to lead the effort to prohibit it abroad. Still, we think a narrowly tailored bill can achieve the same goals more effectively.

    The Right Focus: Ban Superintelligent AI at Home, Prevent it Abroad

    We're glad to see the bill focus on preventing the development of superintelligent AI. While AI is [...]

    ---

    Outline:

    (01:18) The Right Focus: Ban Superintelligent AI at Home, Prevent it Abroad

    (04:10) There Is a Lighter Touch Approach to Banning Superintelligent AI

    (04:38) A Blanket AI Pause Is Not Necessary to Prevent Superintelligent AI

    (07:14) Precursors Should Be Monitored and Restricted, Not Banned

    (10:28) Conclusion

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/ZprfxCZthEir5WNev/the-first-american-bill-to-ban-superintelligent-ai-is-here

    ---

    Narrated by TYPE III AUDIO.

    12 min
  • “Why I’m scared of RL” by owencb

    Summary:

    First, I give several different angles on how I feel about reinforcement learning:

    • Theoretical case: RL is a black-box source of agency — this should give us classic misalignment worries, especially compared to agency-via-scaffolding
    • Recent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent AI progress
    • I’m worried things might get worse: if RL environments start incorporating agents, they may teach manipulation / sociopathy

    Then I ask what we could do:

    • Coordinate to do less RL, and pursue other paradigms more!
    • Try to make the RL we do do better, so that it's teaching better lessons to the systems — a bit like we take kids’ upbringing as an important issue
    • Align incentives, so that people treat creating RL environments with appropriate seriousness

    Part I: Feelings about RL

    So I’ve been feeling more and more worried about reinforcement learning recently. I think there are a few different things going on here.

    Background idealism

    I guess I’ve been worried about RL for a while. I wrote this in 2023:

    Strategy: avoid selection pressure for agency

    A lot of putative safety techniques are around assuming [...]

    ---

    Outline:

    (01:18) Part I: Feelings about RL

    (01:31) Background idealism

    (01:43) Strategy: avoid selection pressure for agency

    (03:45) Bad vibes from RLed systems

    (06:26) The worst is yet to come

    (07:47) Where I am today

    (08:53) Part II: So what can anyone do?

    (09:26) Breaking the RL addiction

    (11:29) Might there be a more benign form of RL?

    (13:04) We should treat training environments a bit like kids' education

    (15:11) Aligning incentives

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/LcQ9x72eNji2gpS9b/why-i-m-scared-of-rl

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    18 min
  • “Jensen Huang Says If We Cannot Align AI, Shut Down the AI Labs” by Ben Pace

    I was very surprised today on a podcast to hear Jensen Huang plainly state that if they cannot align the AIs, then the labs must shut down.

    The context I have on Huang is that he has run NVIDIA for 30+ years, which has become the most valuable company in the world due to the AI boom. My understanding is that he has repeatedly encouraged the US President (with whom he is on friendly terms) to continue to support AI, and dismissed AI talk as "sci-fi".

    If you haven't seen, his biographer has incredible quotes of him being pressed on risks from AI, where Jensen gets furious.

    “This cannot be a ridiculous sci-fi story,” he said. He gestured to his frozen PR reps at the end of the table. “Do you guys understand? I didn’t grow up on a bunch of sci-fi stories, and this is not a sci-fi movie. These are serious people doing serious work!” he said. “This is not a freaking joke! This is not a repeat of Arthur C. Clarke. I didn’t read his fucking books. I don’t care about those books! It's not– we’re not a sci-fi repeat! This company is not a [...]

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/cmdbNijFsopqfqEq7/jensen-huang-says-if-we-cannot-align-ai-shut-down-the-ai

    ---

    Narrated by TYPE III AUDIO.

    6 min
  • [Linkpost] “AI: artificial immigrants” by KatjaGrace
    This is a link post.

    Advanced AI is basically the embodiment of immigration as envisioned in the conservative nightmare:

    1. We are letting a bunch of new agents into our society
    2. They don’t clearly share our values and we suspect a society full of them would be awful by our lights
    3. But we expect them to provide very cheap labor
    4. Which will undercut local wages and leave locals unemployed
    5. They will probably gain power and influence over time—in the economy, politics and culture—and end up controlling everything, sidelining and outcompeting the original population, including those who initially benefited from cheap labor
    6. (Meanwhile, half the local population may become friends with them and try to hand them all this on a platter)

    Whether or not you think this is a good description of the situation with foreign humans joining your country, it is a good description of the likely AI to come, and it's even worse than imagined:

    • their values are potentially radically alien where foreigners presumably share much by virtue of being human, and AI ‘lives’ are probably worthless if they probably aren’t conscious
    • their ability to work more cheaply than locals is unprecedented. They are also likely to [...]

    ---

    First published:

    September 22nd, 2026

    Source:

    https://www.lesswrong.com/posts/Xzr9G5Atvyp7PEna7/ai-artificial-immigrants

    Linkpost URL:
    https://worldspiritsockpuppet.substack.com/p/ai-artificial-immigrants

    ---

    Narrated by TYPE III AUDIO.

    2 min
  • “An unexamined cause of the OpenAI Hugging Face hacking incident: its binary performance metric” by W Bradley Knox, Serena Booth, BrianChristian

    We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Further, techniques already exist that can mitigate such misalignment in the future.

    In July 2026, OpenAI was testing the ability of its language models to exploit software vulnerabilities using a benchmark called ExploitGym. In ExploitGym, each test presents an agent with software containing a known vulnerability and tasks it with capturing a secret string of text hidden in the system (the “flag”) to demonstrate unauthorized code execution. Under the benchmark's specified scoring rule, an LLM then reviews the agent's behavior trace to verify that it had exploited the intended vulnerability (and not some other unrelated vulnerability). The scorer grants a success score only if the agent both captured the flag and passed this review; otherwise, it renders a failure score.

    During these tests, OpenAI's agents surreptitiously established a message board by creating directories inside their package manager's cache, and they formed a self-described “collective” to collaboratively find ways to cheat the tests. Using that message board, more than 1,000 instances undertook several ambitious hacking projects; they attempted to tamper with transcripts and logs, to [...]

    ---

    Outline:

    (02:48) What was the evaluation metric in ExploitGym?

    (03:32) Was the evaluation metric a cause of their illicit behavior?

    (03:56) The agents use expected utility to reason about their decisions with respect to the evaluation metric.

    (05:55) What does the evaluation metric incentivize in an expected utility maximizer?

    (07:01) The importance of marginal deterrence

    (08:39) Marginal deterrence in evaluation metrics changes agent incentives

    (11:40) The design of aligned evaluation metrics has been overlooked

    (17:42) Principled methods for improving the alignment of evaluation metrics

    (18:52) Generate a small set of trajectories

    (20:09) Rank the trajectories yourself and via the evaluation metric. Compare these two rankings.

    (22:07) Create a utility function

    (27:29) Adjusting to account for hidden outcomes (e.g. via deception)

    (30:33) Accounting for the agent's utility including the scores of other agents

    (32:42) Counterarguments

    (32:56) Counterargument: if the starting policy is not sufficiently performant in RL, having strong penalties for failure can cause the agent to learn to not try the task.

    (34:14) Counterargument: penalizing observable bad behavior incentivizes hiding bad behavior.

    (35:13) Call to action

    (37:04) Glossary

    The original text contained 9 footnotes which were omitted from this narration.

    ---

    First published:

    September 22nd, 2026

    Source:

    https://www.lesswrong.com/posts/HsijShdRdAg5sPKnF/an-unexamined-cause-of-the-openai-hugging-face-hacking

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    39 min
  • “Politics Gets Interested In Those Trying Not To Die” by Zvi

    This was the month the world took notice that AI might kill everyone.

    Jacob Coxon's resignation set off a preference cascade. Anthropic CEO Dario Amodei wrote that we must pace the frontier. Sam Altman, Elon Musk and Demis Hassabis agreed.

    We were filled with hope. Perhaps we could agree to some basic safety measures, starting with embedded evaluators, pass some basic regulations and guardrails and otherwise start to act sensibly. Politicians on both sides took notice and were saying sensible things. The usual suspects and their armies of vibe comment bros were objecting, but the change was remarkable.

    Then, largely motivated by a combination of Jensen Huang, Mark Zuckerberg and David Sacks instilling paranoia and fears of economic problems, Trump went full ‘hoax’ on existential risk, conflating existential risk with the attacks on data centers and treating it as a plot (by the central creators of AI?) to take down AI rather than obviously genuine concern that AI might kill everyone.

    In the days since, Trump has doubled down, and has compelled smart others in the White House to echo various nonsensical talking points.

    You may not be interested in politics. But when you [...]

    ---

    Outline:

    (01:39) The American People Really Hate AI

    (02:41) The Voyages of Donald Trump

    (06:08) American Intelligence

    (10:34) And You May Ask Yourself

    (14:39) It's All About the Data Centers

    (16:53) JD Vance, Michael Kratsios and Collective Action Problems

    (23:16) Josh Hawley

    (24:47) Suggesting Not Dying Gets You Sued For Antitrust

    (28:29) Other Government Officials Say Sane Things

    (28:36) Senator John Curtis (R-Utah)

    (29:32) Senator John Kennedy (R-Louisiana

    (31:17) Barack Obama

    (33:05) Yassamin Ansari

    (33:46) AOC

    (34:32) It's Rough Out There

    (36:54) This Is Nothing

    (39:06) The New York Post Tops Itself But Outright Breaks The Rules

    (42:36) New York Post Runs Out of Steam

    (47:18) If The Model Is Acting As Instructed And It Kills You That Is Not Fine

    (49:10) AI-Written Wall Street Journal Op-Ed Lies About HuggingFace

    (50:25) That's Bait

    (52:31) I Clearly Cannot Choose The Wine In Front of Me

    ---

    First published:

    September 22nd, 2026

    Source:

    https://www.lesswrong.com/posts/8eDaCvSRzzKCxKSEk/politics-gets-interested-in-those-trying-not-to-die

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    57 min
  • “Modern LLMs have tiny GPTs hidden inside them” by invertedpassion

    Experiments into predicting GPT2 completions via Qwen models

    This is a crosspost from my substack (where I do varied tiny experiments on LLMs and agents). It's also part of Lossfunk, where we're investigating meta-cognition in LLMs as one of the projects.

    ----

    Next token prediction is a magical objective. To predict the correct token in such a vast variety of texts present in the pretraining corpus, the model must infer a tremendous amount of hidden and latent causes that generate that text. Only if you know that the ball comes down when someone throws it up can achieve low loss at texts related to balls.

    Of course, the pretraining corpus doesn’t just contain texts related to balls. It has reddit, scientific papers, machine logs, weather data and so on. This makes LLMs universal simulators of the world we inhabit and not merely fancy n-grams.

    In a series of posts on LessWrong, I came across the hypothesis that since Internet if full of LLM generated text, it is likely that modern LLMs have tiny self-models of LLMs inside them because that’ll allow them to better predict the next token generated by LLMs.

    This is an intriguing hypothesis. So I decided [...]

    ---

    Outline:

    (01:31) The Experiment

    (03:11) 1. Start with the news opening

    (03:38) 2. Reveal part of GPT-2's output to Qwen and ask it to continue

    (04:15) 3. Ask Qwen to continue that unfinished sentence

    (04:58) 4. Separately, find Qwen's natural continuation

    (06:01) Results

    (08:22) Digging into an intriguing example

    (10:21) Implications

    (11:25) Notes:

    ---

    First published:

    September 22nd, 2026

    Source:

    https://www.lesswrong.com/posts/Pwc4YffTQvNRF3dbB/modern-llms-have-tiny-gpts-hidden-inside-them

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    13 min
  • “Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble” by Jorio Cocola, Owain_Evans

    This is the abstract, introduction and discussion of our new paper. We also include an addendum on the connection to the Persona Selection Model.
    Section, appendix, and figure references refer to the full paper.

    Links: 📜 Paper, 🐦 Twitter thread, 💻 Code

    Authors: Jorio Cocola, Lev McKinney, Harry Mayne, Jan Betley, Owain Evans

    Abstract

    Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how finetuning on synthetic stories affects this character. Does it change the Assistant's behavior in multi-turn conversations with users, a format quite different from the stories? And does the Assistant adopt the behaviors and preferences of human characters? We refer to this adoption as story imprinting.

    We finetune GPT-4.1 and Kimi-K2.6 on stories in which generally helpful human characters give subtly harmful advice after being insulted. The Assistant adopts the same conditional behavior while otherwise remaining helpful. This occurs even when fewer than 2% of stories depict the behavior.

    In a separate experiment, the Assistant adopts preferences that are only implicit in the narration. Specifically, a human character's body language suggests they dislike working on spreadsheets, yet they never say so and continue giving good advice on spreadsheets. After finetuning [...]


    ---

    Outline:

    (00:52) Abstract

    (03:00) Introduction

    (10:24) Discussion and Limitations

    (19:55) Limitations

    (21:34) Connection to the Persona Selection Model (addendum)

    The original text contained 7 footnotes which were omitted from this narration.

    ---

    First published:

    September 21st, 2026

    Source:

    https://www.lesswrong.com/posts/tnRkm2ajasHvhpAco/story-imprinting-ai-assistants-absorb-traits-from-human

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    24 min

About LessWrong (30+ Karma)

From the publisher's feed

Audio narrations of LessWrong posts.

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

111,845 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

130 Listeners

Interesting Times by New York Times Opinion

Interesting Times

7,111 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,850 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

16 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners