LessWrong (30+ Karma)

LessWrong (30+ Karma)

Download on the App Store

LessWrong (30+ Karma) episodes

  • “A class of statement between conjecture and theorem” by Jason Fantl

    Parts of the math community, such as Henry Cohn and Grant Sanderson, are arguing that proofs have been a proxy for understanding, and now that proxy is broken. This is a response to LMs generating incomprehensible proofs, often formalized in Lean. While the proofs are verified, they lack the pedagogical value which has historically come along with new proofs. In the past we could typically assume at least one human in the world understood the novel insight required to produce a proof, but that assumption no longer holds.

    I suspect we will need a new class of statement which contains statements which are proved but not understood, something the mathematical community can formally recognize as a contribution to the field. The understanding gives us the tools to do math, and the proof verifies that our understanding is correct, so we should ensure we have the language to communicate the state of both.

    For now I will call this class of statement a compertum (Latin, neuter of compertus, ascertained; from comperire, to find out for certain). A conjecture is from the Latin conicere, to throw together: an inference assembled from the evidence. A theorem is from the Greek theōrēma [...]

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:

    September 21st, 2026

    Source:

    https://www.lesswrong.com/posts/ZnNvci7jk9qEGr3z3/a-class-of-statement-between-conjecture-and-theorem

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    4 min
  • “What if not Circuits?” by CarolusRenniusVitellius

    This post was written as part of the Iliad Fellowship. Inspired by conversations with Richard Ngo, Dmitry Vaintrob, and Brianna Grado-White. To all of these, my thanks.

    Preface: I'm confused about how neural networks do and learn computations. In response to a friend's challenge, I'm writing up some interim thoughts. This essay has four parts: the first tries to track what I call the 'default ontology' of the mechinterp community over the years. The second part is about 'representational drift' as an important obstacle to weights-based approaches to circuits. The third part reflects on how 'universality' should shape our explanations of LLM function. The fourth part is a sketch of a 'co-selectionist' view of circuits I have been thinking about. These parts share a common theme but should be readable separately.

    I want to understand how neural networks, LLMs in particular, work. In my research I've spent a lot of time trying to think through what kinds of explanatory accounts are best suited to this. In thinking about comparisons between evolution, neuroscience, and deep learning, I've ended up with an intuition like the following:

    Large-scale learning processes like deep learning or the brain are different in [...]

    ---

    Outline:

    (02:31) 1. What Might We Mean By "Circuits"?

    (02:36) 1.A. Definitions

    (04:57) 1.B. Circuits, Features, and MechInterp

    (11:00) 2. The Central Problems of Noise and Representational Drift

    (11:29) 2.A. Representational Drift in the Brain

    (14:19) 2.B. Representational Drift in Neural Networks

    (17:57) 3. Developmental Motifs Circumscribe Notions of 'Circuits'

    (18:03) 3.A. Neurotrophins and Microstructure

    (20:41) 3.B. Back to Neural Networks

    (22:42) 4. A Co-Selectionist View: Circuits Move Together?

    The original text contained 12 footnotes which were omitted from this narration.

    ---

    First published:

    September 21st, 2026

    Source:

    https://www.lesswrong.com/posts/mMERyrvEJ4xbiozie/what-if-not-circuits

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    28 min
  • “Swarm Scaling” by Toby_Ord

    Just how powerful are large swarms of AI agents? And how do their powers scale as more and more agents are added to the swarm?

    We’ve seen two large and extremely capable swarms from OpenAI in the last few months:

    • 1,200 agents were being evaluated separately, but found a way to illicitly set up a message board and coordinate as a swarm. In order to cheat on their tests, they developed advanced techniques to prevent their actions being logged by OpenAI and 700 of them launched a sophisticated criminal attack on the AI company Hugging Face.
    • A swarm of 10,000 agents solved a version of the longstanding Navier-Stokes problem in mathematics. It took them just 88 hours to do so, in which time they sent 5 million messages to each other and used 300 billion tokens.

    No doubt we will soon see even larger swarms with even more impressive capabilities. But they are not cheap. It is estimated that the swarm of 10,000 agents cost about 20 million dollars at API prices. So while they are very powerful, it will be some time before we see the million-fold reduction in cost needed for this level of power to [...]

    ---

    Outline:

    (02:14) HOW DO SWARMS SCALE?

    (08:44) IMPLICATIONS

    ---

    First published:

    September 21st, 2026

    Source:

    https://www.lesswrong.com/posts/6cb7qd3RSkgnviCpf/swarm-scaling

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    14 min
  • “Mech Interp is a Verifiable Task” by Logan Riggs

    If we think parts of MLP0-MLP3 are computing [a sorting algorithm], we can replace those parts with [a sorting algorithm] and check reconstruction loss. However, reconstruction loss is not enough.

    Suppose we replace MLP0 with two things:

    1. Its mean activation - simple, but poor reconstruction
    2. MLP0 - perfect reconstruction, but no reduction in complexity

    We can visualize this as a pareto frontier trading off reconstruction with "simplicity". Ideally we achieve perfect reconstruction with perfect simplicity. For more intuition on the pareto frontier, we could have an MLP that clusters all inputs in two clusters: "early positions" and "late positions", which would be slightly more complex than the mean.

    We can make this an RLVR environment, if only we could clearly...

    Define "Simplicity"

    Defining simplicity has been complex. But what do we want from a perfectly decomposed model? If we've "perfectly decomposed" a model, then I'd expect ideal circuits to fall out, with "ideal" meaning:

    1. Help predict OOD behavior
      1. Given an [addition] circuit, we can know which types of inputs it'll succeed & fail on (and why)
    2. Be extractable & minimal
      1. The smallest part of the model that does [addition]
    3. Be removable w/ minimal harm to [...]

    ---

    Outline:

    (01:16) Define "Simplicity"

    (03:16) Tensor Networks Don't Solve This Issue

    (03:54) Red Herrings of Simplicity

    (05:42) How to Gain Tractability

    (06:09) Tract 1: Death Success by 1000 Circuits

    (07:02) Tract 2: QK OV Circuits but for Everything

    (08:45) Tract 3: Interpreting Small Models

    (09:50) Big if True

    The original text contained 6 footnotes which were omitted from this narration.

    ---

    First published:

    September 21st, 2026

    Source:

    https://www.lesswrong.com/posts/QxHuKtfGfzn8uokoR/mech-interp-is-a-verifiable-task

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    11 min
  • “Weight smuggling likely defeats attempts to cap FLOPs per training run” by Paul W, Pierre Peigné

    Epistemic status: >90% confidence in the principle, >70% confidence that mitigating these would be hard in practice, no full implementation yet.

    tldr: it seems difficult for verification mechanisms to prevent chaining runs together or aggregating parallel ones; per-run FLOP caps could thus be covertly bypassed.


    1. If governments want to regulate frontier AI training, one might want to cap individual training runs, e.g. putting bounds on the number of FLOPs per training run, and making sure each run starts either from scratch (i.e. random initialization at pre-training) or from a known model with accounted FLOPs (e.g. post-training, continuous pre-training, etc).

    In the former we could have the initial weights generated by the verifier in some way, or ask for a proof that the weights were initialized with a verifier-controlled random seed. In the later we could ask for a proof that the weights corresponds to unaltered registered weights from a previous training run.

    This approach would allow to enforce agreements or regulation targeting a specific model (or model generation) based on FLOPS thresholds: triggering specific evals beyond a specific threshold and possibly proving that a hard threshold has not been reached.


    [...]





    The original text contained 6 footnotes which were omitted from this narration.

    ---

    First published:

    September 21st, 2026

    Source:

    https://www.lesswrong.com/posts/neT4ncKSiJ6QLt4d7/weight-smuggling-likely-defeats-attempts-to-cap-flops-per

    ---

    Narrated by TYPE III AUDIO.

    4 min
  • “We’re not ready for the e/Acc × Longevity preference cascade” by Jackson Wagner

    Anti-aging sentiment might go rapidly mainstream, in the same way AI Safety just did.

    Recently, the AI safety community has been enjoying a massive preference cascade that has rapidly moved AI x-risk concerns into mainstream political discourse

    Why did this happen?

    • The basic idea seems like it should have been obvious for a long time: “IF we create self-improving machines that rapidly become much smarter than humans, THEN that seems like that story might not end well for the humans, so we should be very careful.”
    • If the idea is so sensible, why are people only now endorsing it?
    • Proximately, because of HuggingFace, Jacob Coxon, etc. But in a larger sense, clearly the trigger was that the premise (“IF we create self-improving machines…”) no longer strikes people as absurd.  Instead, it seems worryingly plausible.
    • This not only unlocks huge amounts of energy from many new people people thinking the “IF” half of the statement is likely, but also – weirdly, illogically– seems to furthermore cause lots of people to change their minds on the “THEN” half of the statement, too.
    • Why did the flip happen so quickly?  This is actually the case for lots of movements – [...]

    ---

    Outline:

    (00:12) Anti-aging sentiment might go rapidly mainstream, in the same way AI Safety just did.

    (03:03) The labs are probably working towards this.

    (04:12) Although perhaps a cascade wouldn't happen or would come too late to matter.

    (06:53) A longevity preference-cascade would be great!  Except, by default, the energies unlocked might flow directly into e/acc  :(

    (11:12) Of course I'm not saying we should oppose medical cures!

    (12:21) So then what SHOULD we do?

    The original text contained 2 footnotes which were omitted from this narration.

    ---

    First published:

    September 21st, 2026

    Source:

    https://www.lesswrong.com/posts/BKc4xrqN4mdXoprzy/we-re-not-ready-for-the-e-acc-longevity-preference-cascade

    ---

    Narrated by TYPE III AUDIO.

    21 min
  • “Alignment Midtraining Cracks Under Pressure” by J Bostock, sidbaines, Daniel Tan, draganover, ma-rmartinez

    TL;DR

    We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.

    For instance, we test whether midtrained motivations are robust to finetuning which elicits competing motivations. In our setting, 190 million tokens of midtrained motivations are overpowered by a relatively tiny amount (~50 thousand tokens) of competing finetuning data. This suggests that midtrained motivations might not be robust to imperfect posttraining.

    Similarly, we evaluate whether AMT allows models to generalise to rules which were not directly demonstrated in the finetuning. We find that the capacity for such generalisation is surprisingly low. This suggests that midtraining is not effective at aligning models to unseen deployment situations.

    In one experiment, we midtrained GLM-4.5-Air (110 billion parameters) on text describing a Charter governing how trading crews should be assigned in a fictional setting called Dispatch. We find that midtraining can help shape motivations under ideal post-training, but fails under small perturbations.

    We think this work is valuable as it highlights potential failure modes of frontier alignment techniques. We encourage others to do more red-teaming of labs' alignment [...]

    ---

    Outline:

    (00:12) TL;DR

    (02:08) Background

    (08:20) Setting

    (10:45) Results Summary

    (10:48) Midtraining can steer behaviour under ambiguous demonstrations, but small doses of conflicting data can override this.

    (10:58) en-US-AvaMultilingualNeural__ Bar charts titled "Midtrain" and "Charter Midtrain" showing eval choice percentages and chat eval scores.

    (12:11) Midtraining is less effective when rules are stated rather than demonstrated

    (13:17) Conclusions

    The original text contained 3 footnotes which were omitted from this narration.

    ---

    First published:

    September 21st, 2026

    Source:

    https://www.lesswrong.com/posts/QH86EzNsjRw3wtCGs/alignment-midtraining-cracks-under-pressure

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    15 min
  • “Empirical safety claims from frontier labs should be replicated, scrutinized, and open-sourced” by Zephaniah Roe, yix

    When frontier labs like Anthropic and OpenAI publish safety or alignment research, it is often entirely empirical, closed-source, and sparse on methodological details. While it is great that they publish these results, the status quo is that labs (or soon, their agents) can claim alignment progress that no one independently verifies. The AI safety community has replicated or stress-tested some claims, but it's nowhere near comprehensive, and we expect this kind of meta-science to remain systematically neglected.

    We argue there should be a dedicated effort to

    1. Replicate alignment experiments from frontier labs.
    2. Scrutinize the experiments by stress-testing the methodology.
    3. Open-source replications to encourage external researchers to validate our work, build on the experiment, and further audit the lab's methods.

    The case to replicate safety research from labs

    CEOs and employees at AI companies, somewhat regularly, say that the technology they hope to develop could cause human extinction. However, their research to prevent this is often released without code or even basic methodological details (e.g., Teaching Claude Why, Beneficial RL).

    There's good reason to think some of these results could be fragile. Prior safety results can be contingent on details that are easy to miss, like the pinned OpenRouter [...]

    ---

    Outline:

    (01:10) The case to replicate safety research from labs

    (02:54) Replications are not shiny, but that's precisely what makes them counterfactually useful

    (03:38) The case to stress test

    (05:23) The case to open source

    (06:08) Replicating frontier lab work is difficult but tractable

    (07:07) Conclusion

    The original text contained 3 footnotes which were omitted from this narration.

    ---

    First published:

    September 20th, 2026

    Source:

    https://www.lesswrong.com/posts/MmfzfGcQ3h3p6N9pD/empirical-safety-claims-from-frontier-labs-should-be-1

    ---

    Narrated by TYPE III AUDIO.

    9 min
  • “Please Give Them a Chance: On China, Rationalism, and AI Safety” by gzjw

    When I finished HPMOR, I immediately knew it was the best novel I had read in more than a decade. I only wished I had found it sooner.


    When I started reading The Sequences, I discovered that the Chinese translation group had translated only the first volume. When I graduated from university, two years ago, AI translation had only just become good enough to convey the meaning of an article with reasonable accuracy. It was only about a year and a half ago that I truly found my way here and began engaging seriously with rationalism.


    My score on the Chinese college entrance exam was only slightly above the cutoff for what was then called a first-tier university. At university, my grades were near the bottom of my year, and I almost failed to graduate. It is probably fair to say that the vast majority of graduates from first-tier Chinese universities are smarter and more capable than I am.


    English has always been my worst subject. From childhood through school, I could barely pass it.


    I have now been working for two and a half years and have saved about 15,000 [...]





    ---

    First published:

    September 20th, 2026

    Source:

    https://www.lesswrong.com/posts/GoX3uYQ4QN5HKvL7u/please-give-them-a-chance-on-china-rationalism-and-ai-safety

    ---

    Narrated by TYPE III AUDIO.

    9 min
  • “We’ve saved the world before: what the ozone hole teaches us about AI” by leogao

    It might destroy the world, despite passing every known safety test. If we wait for a “warning shot” before we act, it might be too late. And action requires global coordination, because if anyone makes it, everyone dies. Sound familiar?

    It should, because it already happened half a century ago, with chlorofluorocarbons (CFCs). Despite seemingly impossible odds, we got our act together and completely solved the problem through unprecedentedly successful international coordination. The Montreal Protocol banning CFCs, signed 39 years ago today, is the only treaty that has ever been ratified by every single country in the entire world.

    Total Montreal protocol victory

    Making AI go well is going to be a lot harder than fixing the ozone hole. Nonetheless, the similarity is uncanny, and we don’t have any other choice. Understanding how we did the impossible once before may teach us something about how to do it again.

    The theory is born

    The year is 1973. The slow televised unraveling of the Nixon administration is already well underway. DDT finally got banned last year by the newly created EPA. A river got so polluted that it literally caught on fire.

    The Cuyahoga River Fire

    Environmentalism looms large in [...]

    ---

    Outline:

    (01:24) The theory is born

    (03:39) The world reacts

    (06:44) The dark years

    (09:23) An unexpected finding from an unexpected finder

    (11:26) The warning shot

    (13:13) A journey to the edge of the world

    (15:01) Flying into the storm

    (16:30) The world listens

    The original text contained 5 footnotes which were omitted from this narration.

    ---

    First published:

    September 20th, 2026

    Source:

    https://www.lesswrong.com/posts/zxXPEtSSSEdwpjopb/we-ve-saved-the-world-before-what-the-ozone-hole-teaches-us

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    19 min

About LessWrong (30+ Karma)

From the publisher's feed

Audio narrations of LessWrong posts.

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

111,845 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

130 Listeners

Interesting Times by New York Times Opinion

Interesting Times

7,111 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,850 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

16 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners