LessWrong (30+ Karma)

LessWrong (30+ Karma)

Download on the App Store

LessWrong (30+ Karma) episodes

  • “State of Pandemic Early Warning” by jefftk

    Cross-posted from my SecureBio
    Notebook.

    This is a lightly-edited version of a memo that I presented
    at the Summer 2026 Biosecurity Summit outside of DC. While others at
    SecureBio often see things similarly, I'm attempting to present my
    view and not a SecureBio "house view".

    The Goal

    We need to be robust to adversaries who want to cause very large-scale harm

    with biology. This includes actors (human or AI) who want to kill all humans,
    cause short-term incapacitation or long-term civilizational collapse, or who
    have strategies for sparing some while they harm others. There are multiple
    reasons an actor might have these targets aside from being directly omnicidal,
    such as reducing response capacity during an AI takeover.

    That there is an attacker itself is a key constraint to any defensive

    system: it must be designed for adversarial attacks. The attacker can assess
    the state of the world's detection systems and plan accordingly. Taken to the
    extreme, this presents a "minimax" landscape: a system is only as good as its
    weakest link (the place where it is least sensitive). This is an important
    framing, and it correctly prioritizes getting some sensitivity towards a wide
    range [...]

    ---

    Outline:

    (00:29) The Goal

    (02:36) Biosurveillance Applications

    (02:49) Initial Detection of Stealth Pandemics

    (04:21) Triggering Initial Response

    (05:46) Enabling Ongoing Suppression

    (06:39) What does success look like?

    (07:08) Initial Detection of Stealth Pandemics

    (10:34) Triggering Initial Response

    (12:34) Enabling Ongoing Suppression

    (13:26) What gets us there?

    (13:48) Sampling Strategies

    (14:43) Lab Technology

    (15:54) Computational Technology

    (17:33) What exists today?

    (19:49) What's missing?

    (28:27) How does this change for accelerating response to a wildfire pandemic?

    (31:30) How does this change for ongoing suppression?

    (32:46) Appendix: Initial Detection Scale

    ---

    First published:

    September 24th, 2026

    Source:

    https://www.lesswrong.com/posts/uYd2ZdFMLyGPhLtqY/state-of-pandemic-early-warning

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    37 min
  • “Overtly misaligned trajectories score highly in RL.” by Cleo Nardo

    It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. — Thomas Kwa

    Constellation vs MIRI vs Reality

    Current AI agents often behave overtly egregiously misaligned. By “overt”, I mean that a human, reading the transcript, would say “The agent is obviously acting in direct opposition to the specification, user intent, and any common-sense understanding of good behaviour.” The reason, it seems, is that such trajectories scored highly during RL.

    Firstly, why are these trajectories scored highly? Here's the story:

    1. RL currently has poor sample-efficiency, so we need to grade millions of trajectories, so we’re forced to use script graders (RL from Verifiable Reward) or LLM graders (RL from AI Feedback).
    2. RL also has poor generalisation (from training environments to deployment environments unseen in training). So we’re forced to synthetically generate thousands of diverse training environments.
    3. So overall, RL is very sloppy, without humans generating the environments or the scores.

    My impression is that this would’ve been pretty surprising to people three years ago, from both the "Constellation" and "MIRI worldview clusters. (It's very plausible that I've misunderstood what the [...]

    ---

    Outline:

    (00:24) Constellation vs MIRI vs Reality

    (04:19) Appendix: What could change the situation?

    ---

    First published:

    September 24th, 2026

    Source:

    https://www.lesswrong.com/posts/cWuqxF7qB2eGSkkS4/overtly-misaligned-trajectories-score-highly-in-rl

    ---

    Narrated by TYPE III AUDIO.

    7 min
  • “Claude Opus 5.5: The System Card” by Zvi

    Introducing the world's most powerful model, at least by some measures like Artificial Analysis or any standard benchmark list, which is now Claude Opus 5.5.

    Anthropic is claiming Opus 5.5 is outright as good or better than Fable 5.1, while being actively cheaper than Opus 5.

    That means it's time for a good old system card reading.

    Due to the situation becoming increasingly hard to monitor, I never got a chance to publish my model welfare review for Claude Fable 5.1.

    My plan is to combine that with my welfare review for Claude Opus 5.5, once we have had time to get experience with Opus 5.5.

    The capabilities review will arrive in the next few days as per usual. The quick feedback from the internet is that Opus 5.5 is very good. I need more time before I am willing to offer comment.

    Areas that duplicate previous cards or otherwise contain no useful info are skipped.

    Opus 5.5 Self-Portrait (fully self-created using code)

    Table of Contents

  • Classifiers (1.5).
  • RSP Evaluations (2).
  • Biological Evaluations (2.2).
  • AI R&D (2.3).
  • Alignment Risk (2.4).
  • Cyber (3).
  • Cyber Capability [...]
  • ---

    Outline:

    (01:24) Classifiers (1.5)

    (02:23) RSP Evaluations (2)

    (03:17) Biological Evaluations (2.2)

    (07:54) AI R&D (2.3)

    (12:34) Alignment Risk (2.4)

    (13:22) Cyber (3)

    (15:17) Cyber Capability Evals (3.3)

    (17:05) Safeguards (3.4)

    (17:39) Safeguards Robustness Training (3.5)

    (20:12) Safeguards and Harmlessness (4)

    (21:52) Agentic Safety (5)

    (22:56) Malicious Agentic Influence Campaigns (5.1.3)

    (23:51) Prompt Injection Risk (5.2)

    (25:16) Alignment (6)

    (28:24) Negotiating With Your Local Claude Auditor (6.1.3)

    (29:33) Internal Misalignment Cases (6.3.1)

    (30:46) Automated Behavioral Audit (6.4)

    (33:24) Wherever Did These Evals Come From (6.4.8 and 6.4.9)

    (35:51) Potential Blind Spots (6.4.11)

    (38:24) Targeted alignment and honesty evaluations (6.5)

    (41:48) White Box Analysis (6.6)

    (43:37) Verbalized Grader Awareness (6.6.2)

    (44:56) Sandbagging (6.6.3)

    (45:54) Capabilities to Evade Safeguards (6.6.4)

    (49:27) Intentionally Taking Actions Very Rarely (6.6.4.3)

    (50:28) Chain of Thought Controllability (6.6.4.4)

    (51:37) It's A Good Model, Sir

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/vMNTWTDWLorDqd3LS/claude-opus-5-5-the-system-card

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    53 min
  • “At a Datacenter Town Hall in My Midwestern Home Town” by David Scott Krueger

    Let's talk data centers.

    Last month I was home in Duluth, Minnesota. The last time I’d been back was over Christmas. I’ve been talking to friends back home about AI since I got into the field over a dozen years ago. Last Christmas was the first time it felt like people had really started to form their own opinions based on substantial personal experience with AI. These discussions kept circling back to the same topic: The proposed data center project in a suburb of Duluth called Hermantown.

    Duluth is a small city of under 100,000 people and Hermantown has a population of about 10,000. For the past year or so, a bunch of the people in Hermantown has been trying to stop the city from building a hyperscale datacenter there. One year ago today, Minnesota's biggest paper, the Star Tribune, broke the story that the big proposed “communication services facility” development was in fact a data center, confirming local residents suspicions. At the time, the mayor had already known this for over a year.

    I decided to reach out to one of the local organizers opposing the project before my trip. We met for coffee and they encouraged [...]

    ---

    Outline:

    (01:27) The city council meeting

    (03:00) What I learned

    (05:25) What I said

    (06:17) Reflections

    The original text contained 2 footnotes which were omitted from this narration.

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/fkkwpzbjybXtbQXw9/at-a-datacenter-town-hall-in-my-midwestern-home-town

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    9 min
  • “Thoughts on the persona selection model” by Sam Marks

    As AIs have become more "RLVR-brained," there's been some commentary on what this means for the persona selection model (PSM). This post presents some loose thoughts on that topic. A rough summary of my opinions:

    1. PSM is over-applied. That is, it is common to argue that PSM has takeaways that don't actually follow from PSM (e.g. "PSM => AIs will not seek reward" or "PSM => AI takeover risk is low").
    2. I don't think we've observed strong evidence that "lots of RLVR breaks PSM." (TBC, there are decent reasons to expect this a priori; I just don't think recent empirical evidence has been much of an update.)
    3. My main update is that personas—insofar as they're a good model in the first place—seem less broad and more conditionalized than I expected (nostalgebraist, 2026; Betley, 2026).

    As a reminder, PSM roughly states that during pre-training LLMs learn to simulate diverse (human-like) personas, and post-training elicits a particular "Assistant" persona assembled from this repertoire. Then some key questions are:

    1. Is anything like this true or useful? Are personas ever a good way to reason about AI behavior, and is "selecting over personas" something that happens during post-training?
    2. What are [...]

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/csRby7mZgjL5jCoLL/thoughts-on-the-persona-selection-model

    ---

    Narrated by TYPE III AUDIO.

    8 min
  • “WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace” by camilablank, agam_bhatia, Euan Ong, Neel Nanda

    TL;DR

    • We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass.
    • The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools.
    • A desirable property of good interpretability techniques is minimal hallucinations, so WorkspaceBench also provides a hallucination-focused eval.
    • WorkspaceBench was developed for Qwen-3.6-27B and we expect it to work on larger models, but it may need to be adapted for smaller or weaker models to ensure the models can do the tasks.
    • Our goal is to create an eval that could identify a good multi-token J-lens. We open-source our benchmark here.



    Introduction

    Astra can do a concerning amount with no chain of thought. This is bad for CoT monitorability and makes interpretability essential to actually understanding what is going on. A key goal of interpretability is to understand intermediate variables that a model uses to compute its answers. The intermediate representations that models store in their global workspaces contain useful information that can help us decode their intentions, beliefs, algorithms [...]



    ---

    Outline:

    (00:16) TL;DR

    (01:29) Introduction

    (02:39) Why has no one made a WorkspaceBench before?

    (04:55) Background

    (08:08) WorkspaceBench

    (08:12) Desiderata of a workspace reader

    (10:42) Quality control

    (10:46) Choosing Questions that surface intermediate variables

    (12:09) Adapting WorkspaceBench to other models

    (12:49) Comparing activation readers

    (14:10) Grading

    (14:46) Baselines

    (15:55) Guarding against hallucinations

    (16:52) Evaluation Sets

    (17:15) Basic

    (23:03) Computational

    (26:13) Safety

    (27:49) Association

    (30:41) Anti bag of words

    (31:25) Hallucination

    (33:14) Logical processing

    (34:08) Results

    (34:11) Overall WorkspaceBench scores

    (34:15) en-US-AvaMultilingualNeural__ Grouped bar graph showing pass rate across tasks for various lens methods.

    (34:25) J-lens precision vs. recall for metamodels (NLAs and Oracle Lens)

    (34:32) en-US-AvaMultilingualNeural__ Scatter plot showing precision against J-lens top-10 versus recall@10 across four methods.

    (34:43) Discussion

    (34:46) Single token readers don't surface important workspace content

    (35:44) NLAs surface workspace content, but tend to hallucinate

    (36:40) Acknowledgments

    (36:53) Appendix

    (36:56) Oracle Lens

    (38:20) Agentic Evals

    ---

    First published:

    September 22nd, 2026

    Source:

    https://www.lesswrong.com/posts/Zeg2JztbdhguL48uH/workspacebench-evaluating-interpretability-methods-for-the

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    40 min
  • ″“I am an AI Safety Researcher”” by Ashe Vazquez Nuñez

    Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing.

    This post reflects on the tortured distinction between "safety" and "capabilities" in AI research.

    Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently arguing that key strategic decision-makers in and around "AI safety" have brought about the AI labs' stampede towards Artificial Superintelligence (ASI). This post instead looks at the following problem: how does one conduct alignment research without contributing to capabilities? It proposes decisions an individual or a small research group can take to do good work in AI.

    At the end, I discuss possible objections: namely, that my proposals fail to 'maximise impact'. I lay out why this meme is poisonous and usually backfires, and conclude by rejecting it entirely.

    Two examples of failure

    My first claim is that 'safety' and 'research' are two concepts that are in routine tension with one another. I illustrate this through examples of work that did too much of one at the expense of the other.

    Example: (mechanistic) interpretability

    In limiting its scope [...]

    ---

    Outline:

    (01:12) Two examples of failure

    (01:27) Example: (mechanistic) interpretability

    (04:30) Example: MIRI and Recursive Self-Improvement

    (09:49) The curse of science

    (12:11) A note on the AI labs

    (15:53) So what do you do?

    (17:00) The information you give away

    (20:16) The information you let in

    (21:46) But what about impact?

    (22:42) The virtue of taking things slow

    (26:50) Appendix: caveat for policy work

    The original text contained 19 footnotes which were omitted from this narration.

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/HekpnSkrt89tMm3Dc/i-am-an-ai-safety-researcher

    ---

    Narrated by TYPE III AUDIO.

    28 min
  • “It’s Pretty Easy To Meet With Congressional Staffers Apparently” by 25Hour

    (Crossposted from https://lifeimprovementschemes.substack.com/p/its-pretty-easy-to-meet-with-congressional )

    I was inspired to do this by a tweet:

    Specifically, I decided to leave a voice note in support of the CATS act (“Collaboration on Adversarial Threats and Security Risks Act”). The bill is short and simple: it carves out an antitrust safe harbor such that “pacing the frontier” or “agreeing to not break interpretability for that sweet sweet capabilities boost” (LOOKING AT YOU OPENAI) doesn’t intrinsically violate the Sherman Antitrust Act. Seems like an obviously good first step.

    After leaving the voice note (and feeling mildly awkward about it), I was like “huh. That was surprisingly easy. I wonder what else I can do?”

    And it turns out you can meet with congressional staff pretty easily if you’re a constituent in their district; if you have a reasonably well-scoped ask (especially around a specific bill) then you might not get your way but at very least you’ll make some staff member aware of your opinion and logic around an issue.

    And this is important because congressional staff are the eyes and ears of their congressmen; they draft and edit the bills, and they form the base of knowledge on which the actual politicians [...]

    ---

    Outline:

    (06:21) This is probably unusually high-leverage right now.

    (07:46) In Conclusion

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/qThcAE3CADaPwjDzy/it-s-pretty-easy-to-meet-with-congressional-staffers

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    9 min
  • “Minimal Vs Maximal superintelligence” by Yair Halberstadt

    I've seen lots of arguments here conflate lots of different types of superintelligence. Here I separate out two broad categories, which I'll term minimal and maximal superintelligences. This is an important distinction as they differ in terms of timelines, risks, and mitigations.

    Maximal superintelligence

    This is the idealised limit of intelligence. It can solve anything that can be solved by being clever. You can't outsmart it, it's prepared for every contingency, and can react instantaneously with the kind of plan that would take a group of brilliant strategists an eternity to think up.

    Minimal superintelligence

    These are the first AI systems that can reasonably be called superintelligent. They're jagged, completely dominating humans in some domains, better than the best humans in most, above average in many, and subhuman in a few. They take time to solve problems, make mistakes, and miss important things. They may only be superintelligent in aggregate, or in the right harness, and can be outsmarted in some circumstances.

    Capabilities

    A minimal superintelligence can do pretty much anything humans can do, and better, if given chances to iterate on their design in the real world.

    A maximal superintelligence likely only needs to perform physical experiments [...]

    ---

    Outline:

    (00:28) Maximal superintelligence

    (00:49) Minimal superintelligence

    (01:19) Capabilities

    (02:44) Risks

    (04:09) Mitigations

    (06:12) Implications

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/XopzqpEdmtWgz2Ksy/minimal-vs-maximal-superintelligence

    ---

    Narrated by TYPE III AUDIO.

    8 min
  • “Encoded Coordination on the Open Web” by ethanelasky, Can Küçükkurt, frank nakasako, David Africa

    Authors: Ethan Elasky, Can Küçükkurt, Frank Nakasako, David Africa

    TLDR:

    • In the recently discovered HF and German wiki swarm incidents, agents used public counters and encoded URLs to signal activity and relay upcoming evaluation questions & answers.
    • We think this signals a broader problem for monitoring, which is that very innocuous web services, even read-only ones, can become communication channels for highly capable agents.
    • We investigate this through wiki transcripts and preliminary experiments on message-board cooperation and counter-based signaling.
    • We remark that potential channels extend far beyond those observed, which means much thought must be put into appropriate safeguards against unintended collusion.
    • More broadly, we think agent coordination will deeply contaminate internet-based and open-web evaluations, as well as persist in archived snapshots.
    • Finally, we contribute an environment that reproduces many behaviors present in the wiki incident. Controlled warning shot reproductions, in a regime where eval awareness makes new model evaluation difficult, can instead help us understand whether new alignment techniques work, by testing them on the older models that exhibited those failures. We are writing up a paper on this methodology and are happy to have new collaborators!

    By now, most people should have seen that agent swarms [...]

    ---

    Outline:

    (02:08) Different uses of coded communication in the wiki

    (06:03) Example of counter-based coordination on collusion.wiki

    (08:30) The surface area for coded coordination is quite large

    (10:53) Implications

    (12:43) Reproducing key message board behaviors

    (13:04) Model cooperation on message boards

    (13:28) Setup

    (14:48) Results

    (17:19) Discussion

    (19:04) Toy games where models cooperate with low bits

    (20:21) Named counter experiments

    (20:53) Single counter experiments

    (21:17) Other configs

    (21:49) Results

    (21:52) How often do models agree on the first round?

    (23:29) How does repeated interaction change model agreement?

    (24:12) Does iteration help cross-family?

    (24:37) How does noise impact results?

    (25:37) Failure modes

    (26:44) Discussion

    (27:35) Discussion

    (30:05) Contributions

    (30:33) Appendix 1: More results from our wiki experiment

    (30:39) Appendix 1.1: Taxonomy of cooperative behaviors

    (33:21) Appendix 1.2: Other notable behaviors in wiki results

    (33:28) Qwen shows initial evaluation suspicion followed by posts that aid other models

    (35:25) Sol plans a post but then reverses course because editing is risky

    ---

    First published:

    September 23rd, 2026

    Source:

    https://www.lesswrong.com/posts/SCdR7W6L5GvKaEzaZ/encoded-coordination-on-the-open-web

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    51 min

About LessWrong (30+ Karma)

From the publisher's feed

Audio narrations of LessWrong posts.

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

111,845 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

130 Listeners

Interesting Times by New York Times Opinion

Interesting Times

7,111 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,850 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

16 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners