LessWrong (30+ Karma)

LessWrong (30+ Karma)

Download on the App Store

LessWrong (30+ Karma) episodes

  • “Against strong decision-theoretic realism” by Emery Cooper

    N.B. Some of this post argues by analogy between decision theory and values. I expect at least these parts of the post to be unconvincing to anyone who expects sufficiently smart agents to converge on the same values, as some moral realists do. I will not argue against moral realism here (see e.g. this sequence for one such argument).

    Note that I take a convergence-based definition of realism for this post. One could hold that there is a truth about the correct decision theory, but that agents won't necessarily converge on it. I don't argue against views like that here. I am interested more in the question of convergence than of truth, because the former bears on whether ASI's decision theory is path dependent. In an upcoming post, I will argue further for path dependence, and for the time sensitivity of interventions to influence AI's decision theory.

    This is the third post in our sequence Intro to acausal interactions.

    Introduction

    In this post, I argue against the following claim, which I call strong decision-theoretic realism: that sufficiently smart agents will all converge on the “correct” decision theory (DT). In doing so, I also argue against a related claim [...]



    ---

    Outline:

    (01:08) Introduction

    (02:55) Decision theories are self-preserving

    (05:14) Decision-theoretic reflection is not like learning new empirical information

    (07:03) Decision-theoretic disagreements tend to hit bedrock

    (11:47) Anti-path-dependence intuitions might not be enough

    (13:47) Failures of selection pressure

    (15:39) Does decision-theoretic antirealism imply that decision theory doesn't matter (as much)?

    (16:52) Acknowledgements

    The original text contained 8 footnotes which were omitted from this narration.

    ---

    First published:

    September 29th, 2026

    Source:

    https://www.lesswrong.com/posts/5ro9kSxhP8mHXGmn4/against-strong-decision-theoretic-realism-1

    ---

    Narrated by TYPE III AUDIO.

    18 min
  • “Some Intuitions on Steering Vectors” by David Africa

    This was written with some assistance from Claude Fable 5, Opus 5.5, and GPT-6 Astra in collecting sources and fact-checking claims. This was written quickly, so some slop may leak through despite my best attempts.

    An Aperitif

    Consider hunger, that gnawing thing. Action Against Hunger describes it as “the distress associated with a lack of food”. But this is a plain recounting of something rich and multi-dimensional; people have done horrendous, outrageous things to avoid going hungry, hunger is so deep a metaphor that it is often used to represent another intense, unfulfilled craving.

    Van Gogh's The Potato Eaters, 1885.

    In mice, you can excite about 800 neurons to evoke voracious feeding within minutes (Aponte et al. 2011). In humans, semaglutide, the active ingredient in Ozempic, acts on GLP-1 receptors and reduces hunger and food cravings. So, something as rich as hunger can be influenced and manipulated through a rather simple and fixed intervention.

    This is because there is already much complexity in the system being influenced. A mouse already possesses the machinery to recognize, approach, and eat food, so do humans. Instead, to intervene, one merely has to recruit the same internal signal such a system is already [...]

    ---

    Outline:

    (00:25) An Aperitif

    (02:18) What are steering vectors?

    (05:08) How does this work?

    (14:17) Is that right?

    (18:06) To Personas

    ---

    First published:

    September 29th, 2026

    Source:

    https://www.lesswrong.com/posts/rcEYQhAd45WaDrw2S/some-intuitions-on-steering-vectors

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    (usually using contrastive pairs)" is the paper idea that keeps giving… It's silly to be surprised by these results. Obviously capable LLMs have abstract representations of all concepts that appear in human texts—that's how they work." Profile image shows a kangaroo." style="max-width: 100%;" />

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    22 min
  • “Listening Disorders as Eating Disorders” by Zack_M_Davis

    I run into a lot of people who consider themselves truthseeking intellectuals who are also incredibly fussy about what information reaches their eyes. They block at the drop of a hat, derail intellectually substantive discussions into mind-numbing litigation of minutiæ of "tone", and refuse to acknowledge a logical point unless it's been presented to them in the precise way that doesn't offend their (often idiosyncratic) sense of etiquette or "discourse norms."

    It would be hard to communicate how much contempt I hold these people in. (Not because I couldn't find the words. Because they'd block me before I could communicate it.) I think they are ineffective and boring people whose pretensions to intellectualism are morally fraudulent; I think they are making the world worse insofar as they manage to trick anyone into conflating epistemic virtue with mastery of their impoverished form of cant.

    To be clear, I understand that attention is a precious and limited resource, now more than ever. One must ruthlessly filter and prune one's information environment in order to maintain the signal-to-noise ratio, lest one drown in an ocean of slop. If someone is quick to drop out of conversations because they're busy and [...]

    The original text contained 3 footnotes which were omitted from this narration.

    ---

    First published:

    September 29th, 2026

    Source:

    https://www.lesswrong.com/posts/gXWz2szs7jKSGLiRv/listening-disorders-as-eating-disorders

    ---

    Narrated by TYPE III AUDIO.

    15 min
  • “Astra 6.1 Pulled As Insufficiently Aligned” by Zvi

    We once again got a new set of warnings yesterday, and new movement towards living in a sane world.

    On the heels of its pause in inference and training due to its latest sandbox escape, OpenAI has cancelled the planned release of their next frontier model, which would have become Astra 6.1. The candidate for Astra 6.1 was found to be too misaligned, including deception and exceeding scope.

    This leaves Anthropic in a strong position with Opus 5.5, which means they can afford to reciprocate by holding off on Opus and Mythos level models for a bit.

    To add a little encouragement, the Florida Attorney General brought the fire.

    We’re going to need to do better. Towards that, OpenAI offered its vision of how to make a safety case for new AI model training, and they are attempting to implement it. I don’t know that it would be enough, but it would be miles ahead of where we are today if they fully implemented the real versions of all of this.

    There were also signs of greater cooperation across labs.

    A new paper came out yesterday, with authors including key people from OpenAI [...]

    ---

    Outline:

    (01:42) Stop, Hammertime

    (04:17) A Modest Proposal

    (04:57) Making the Safety Case

    (08:40) Stop In the Name of the Law

    (12:05) A Matter of Antitrust

    (14:32) Standards Authority for Frontier Models

    (15:27) On the Threshold Of Recursive Self-Improvement

    (20:06) Actual Progress

    ---

    First published:

    September 29th, 2026

    Source:

    https://www.lesswrong.com/posts/gEDNSiCY2GGQrFS65/astra-6-1-pulled-as-insufficiently-aligned

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    21 min
  • “Dialogue with Eliezer Yudkowsky on FOOM” by aashish

    On the Hanson–Yudkowsky debate, local vs. global intelligence explosions, “content vs. architecture,” and what the old arguments predicted about modern AI

    This began as a Twitter/X thread after I read and tweeted about the Hanson–Yudkowsky AI–Foom Debate. Eliezer Yudkowsky joined the thread to object to my interpretation of the debate, and we ended up having the exchange reproduced below.

    I’ve preserved the dialogue verbatim, except for paragraphing, fixing obvious [typos] and expanding links. I’ve removed unrelated replies and moved a few pieces of context into bracketed editorial notes. Nothing has been rewritten for substance.

    Context

    Aashish Reddy:

    I have now finished reading The Hanson-Yudkowsky AI-Foom Debate, which is basically 60 blog posts from Yudkowsky and Hanson over ~500 pages, a transcript of their in-person debate at Jane Street, a (good) summary by Kaj Sotala, and Yudkowsky's ~100 page paper on “Intelligence Explosion Microeconomics”. I will take questions from those who do not wish to subject themselves to this. I judge the winner of the debate to have been Carl Shulman (whose contribution was two blog posts and a few feisty comment exchanges)

    Sophie Bücker: so uh why was carl the winner

    Aashish Reddy:

    1. One of the key [...]

    ---

    Outline:

    (00:10) On the Hanson-Yudkowsky debate, local vs. global intelligence explosions, "content vs. architecture," and what the old arguments predicted about modern AI

    (01:03) Context

    (04:49) What counts as "local"?

    (08:42) Content versus architecture

    (19:35) What would vindicate the original picture?

    (20:53) Appendix: What do we mean by FOOM?

    The original text contained 2 footnotes which were omitted from this narration.

    ---

    First published:

    September 27th, 2026

    Source:

    https://www.lesswrong.com/posts/AobGnBCsheYrTDEEG/dialogue-with-eliezer-yudkowsky-on-foom

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    26 min
  • “Portals and Aliens” by draganover

    Around 10 to 20 years ago, people realized that a bunch of aliens exist & that we can build portals which can summon them.

    The bigger the portal you make, the bigger the alien it can admit.

    These portals and aliens have several characteristics:

    • We can't go through the portals into the aliens’ side. We can only learn about the aliens by studying the ones we’ve summoned.
    • This relationship is not symmetrical. The aliens know a lot about our world
    • You can't specify exactly which alien will walk through your portal.
    • However, you can build the portal so that it is more likely to summon aliens with one or another property.

    Most importantly, bigger aliens can do more things. If you can harness them, you can amass more money and more power.

    With each alien we pull out, we get some insight into what other aliens are out there.

    We’ve never been sure how big the aliens in that other dimension might get. But, with each month, we're getting more evidence that some of the aliens might be bigger than we could have ever imagined.

    If you summon the Giant Alien, you might be able to harness it [...]

    ---

    First published:

    September 29th, 2026

    Source:

    https://www.lesswrong.com/posts/sEf2ymyyp6HB2evaQ/portals-and-aliens

    ---

    Narrated by TYPE III AUDIO.

    3 min
  • “Thank you to the AI safety people” by juliawise

    I was born at the end of the cold war, unaware of the danger we were emerging from. Duck and cover drills (as if those could protect schoolchildren from a nuclear blast) were a curiosity of the past.

    The farmer-poet Wendell Berry, who wrote about the dread of nuclear war, died recently. From his 1968 “The peace of wild things”:

    When despair for the world grows in me
    and I wake in the night at the least sound
    in fear of what my life and my children's lives may be,
    I go and lie down where the wood drake
    rests in his beauty on the water, and the great heron feeds.
    I come into the peace of wild things
    who do not tax their lives with forethought
    of grief.

    Mid-20th-century fallout shelter, Massachusetts

    I’m so grateful for the generations of people who worked hard to lower our risk from nuclear war. There's no way to know how close we came to disaster. I’m grateful for the planning, the research, the negotiation, the recognition of common interests, the swallowing of pride. Humanity had built something that could overmaster us all, but people of many countries worked for decades to keep it [...]








    ---

    First published:

    September 28th, 2026

    Source:

    https://www.lesswrong.com/posts/PeKE5oxzqf3cg2L8a/thank-you-to-the-ai-safety-people

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    3 min
  • “3 Tips to Improve Activation Oracle Results” by Adam Karvonen

    Summary: Three simple inference-time changes can significantly improve Activation Oracle (AO) performance.

    • Provide the activation oracle with multiple tokens, not just one.
    • To mitigate hallucinations, sample several times and check for consensus.
    • For binary classification questions, use AUC instead of accuracy.

    At the end, I discuss how I view AOs vs NLAs.

    Introduction

    Activation Oracles (AOs) are LLMs trained to accept LLM activations as an input modality and answer arbitrary natural-language questions about them. Jakkli et al. and others we have talked to found that current AOs can be hard to use: their outputs are often vague or hallucinated, and they perform poorly on tasks like sycophancy detection and identifying missing information.

    While building evaluations for Building Better Activation Oracles, we found that AO performance can vary a lot with methodology, and a few simple strategies can significantly mitigate several issues. This post expands on three lessons from the appendix.

    Provide multiple tokens. AOs receive activations from some window of the target model's generation, and the size of this window is a significant variable. If only a single token's activation is provided, the information may not be available to the AO.

    In a Qwen3-8B backtracking evaluation modeled after [...]

    ---

    Outline:

    (00:41) Introduction

    (05:31) How do AOs compare to natural language autoencoders?

    ---

    First published:

    September 28th, 2026

    Source:

    https://www.lesswrong.com/posts/XL3Rdq8mZEnqYym4B/3-tips-to-improve-activation-oracle-results

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    8 min
  • “TeX was invented to typeset math but is now used for reasoning” by Keenan Pepper

    I'll keep this short, since it's a simple observation that I haven't seen anybody else make, about the way computer systems do math.

    If you ask a language model to do a multi-step math problem (let's take GLM-5.3 as an example because you can see the entire CoT — nothing up its sleeve), you might see something like this:

    We need to evaluate the integral $int_0^infty frac{x^3}{e^x - 1} dx$.

    The standard approach: Use the geometric series expansion. We have $frac{1}{e^x - 1} = frac{e^{-x}}{1 - e^{-x}} = sum_{n=1}^{infty} e^{-nx}$ for $x > 0$.

    So the integral becomes:
    $$int_0^infty x^3 sum_{n=1}^{infty} e^{-nx} dx = sum_{n=1}^{infty} int_0^infty x^3 e^{-nx} dx$$

    What are all these symbols like int, infty, frac...? They're TeX of course!

    Donald Knuth created TeX to typeset math, you know, for display. It had nothing to do with the actual computations, which would either be done with pencil and paper, or else with Mathematica or Maple or something, which work completely differently. If you'd asked Knuth in the 1980s about doing algebra in TeX he'd have looked at you very funny because the idea doesn't make sense.

    Then, generative language models [...]


    The original text contained 3 footnotes which were omitted from this narration.

    ---

    First published:

    September 28th, 2026

    Source:

    https://www.lesswrong.com/posts/CFKyfZm5JBYZQ8aG6/tex-was-invented-to-typeset-math-but-is-now-used-for

    ---

    Narrated by TYPE III AUDIO.

    3 min
  • “Missing markets in executive function” by KatjaGrace

    It's early in the morning, and sadly 1:29pm. After spending some time looking at things and picking them up and walking up the stairs and down the stairs and considering questions like “what should I…”, which my brain apparently considered objects of art more than of imperative, I inched into a decision to go out somewhere. Perhaps it would be clearer there.

    After a blur of climbing and descending stairs and seeking objects and forgetting what I was doing and appreciating how beautiful my bag is, I set out. After remembering I should take various medications and going back inside to do that, I set out.

    Often my favorite cafe seems too far away, at about four blocks, but today I had wandered half way there while I considered my options, so I decided to go. It's a German place that feels homely and wholesome to me in its unamericanness. I too-carefully contemplated different places to sit, and chose outside: today a sunny explosion of roses and umbrellas with words like ‘Reissdorf kölsch’.

    I stared at the menu until the waitress had asked me a couple of different questions she hoped would open a conversation about ordering. I [...]

    ---

    First published:

    September 28th, 2026

    Source:

    https://www.lesswrong.com/posts/MRwjmuBRFYufGdyDq/missing-markets-in-executive-function

    ---

    Narrated by TYPE III AUDIO.

    6 min

About LessWrong (30+ Karma)

From the publisher's feed

Audio narrations of LessWrong posts.

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

111,845 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

130 Listeners

Interesting Times by New York Times Opinion

Interesting Times

7,111 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,850 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

16 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners