LessWrong (30+ Karma)

LessWrong (30+ Karma)

Download on the App Store

LessWrong (30+ Karma) episodes

  • “Claude Opus 5.5 Should Raise Your Ambitions” by Zvi

    When it comes to making things, or doing most things in general, Fable 5.1 and especially GPT-6 Astra raised my ambition level. They should have raised yours, too.

    Claude Opus 5.5 should raise your ambition levels again. It just works, and it persists, like Astra does. It does the things. And it is highly pleasant to talk to, and its writing is pleasant to read, while you are at it. The game has been changed, again.

    Feedback is almost universally positive. Claude was never gone, but also is so back.

    The benchmarks are excellent, but ignore the benchmarks. Be ambitious. Go out and do things. Get curious. Have more interesting conversations.

    If one of those things is Pacing the Frontier or otherwise ensuring that AI does not kill everyone, leaving us to enjoy our bounty? That's even better.

    By Claude Opus 5.5, for this post

    The Official Pitch

    The pitch is Fable-5.1-level performance at lower Opus-level price. Good pitch.

    We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than [...]

    ---

    Outline:

    (01:22) The Official Pitch

    (04:38) Our Price Cheap

    (06:18) Official Benchmarks

    (07:26) Other People's Benchmarks

    (10:54) Claude Classifies

    (12:27) The System Prompt

    (12:34) Reaction Rules

    (13:07) Vision In 3D

    (15:29) Claude Creates

    (18:11) Claude Composes

    (18:53) Positive Reactions

    (24:50) Good Talk

    (27:03) On Writing

    (30:30) Big Model Smell

    (32:30) Check Your Work

    (32:56) Negative Reactions

    (34:02) Not So Fast

    (34:55) Some People Need Practical Advice

    ---

    First published:

    September 26th, 2026

    Source:

    https://www.lesswrong.com/posts/rtPiip9igy3QvxYdM/claude-opus-5-5-should-raise-your-ambitions

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    37 min
  • “Poverty in the midst of abundance: AI will make goods cheaper, but your labor will get cheaper faster” by cousin_it

    Very simple idea, but I thought it'd be worth making a reference post on this.

    Some people are saying AI will make all goods cheaper, so you'll be able to afford a nice life by working. Without any redistribution, just by market mechanisms. These people are wrong.

    AI will lower the price of goods you need to survive, and also the price of your labor. The question is which will get cheaper faster. Let's use energy cost as a proxy. A day's worth of labor equivalent to yours can be done by AI for just a few cents in electricity. But feeding you with e.g. apples for a day will cost more energy than that, because growing apples is harder to energy-optimize than generating tokens. So selling your labor at market price will leave you unable to afford apples.

    This means a future with economic AI might look like "poverty in the midst of abundance". All goods are cheap, and tokens are cheap, but somehow you can't find a job paying even that much.

    Maybe the problem can be solved by redistribution, or by everyone having investments, or something else. That's a bigger discussion. In this post I just [...]

    ---

    First published:

    September 26th, 2026

    Source:

    https://www.lesswrong.com/posts/eLXTcJfkheLbqZXHa/poverty-in-the-midst-of-abundance-ai-will-make-goods-cheaper

    ---

    Narrated by TYPE III AUDIO.

    2 min
  • “Plan R: AI Safety by ASICs” by Roko

    Much of the civilization-scale risk we are seeing in AI in 2026 comes from the following combination: we created a single institution (the "Frontier AI Company") that has two properties:

    A. It is set up to create very powerful and/or self-replicating entities that may exceed the capabilities of the entirety of the rest of civilization and come with extraordinary risks

    B. It gets to own an unbounded financial claim on the resulting surplus

    All the technical stuff about AI, AI alignment, etc can be rolled up into point (A) above. My claim is that having point (A) on its own, without point (B) is probably okay. Nuclear technology and bioweapon technology both approximate (A) and they are mostly okay because without (B), there isn't an incentive for people controlling them to push their luck on safety.

    But with Frontier AI Companies, we mixed the two.

    The key claim of this post is that we can probably get rid of most of AI risk without doing anything other than separating out the bookkeeping, physical footprint and institutions so that there is no single org with both properties. And with a little help from ASICs, maybe we [...]

    ---

    First published:

    September 25th, 2026

    Source:

    https://www.lesswrong.com/posts/n8u3BfqFoGh4jnzpo/plan-r-ai-safety-by-asics

    ---

    Narrated by TYPE III AUDIO.

    10 min
  • “Evidence about risk should be transparent” by Ajeya Cotra

    All views are my own and do not represent my employer.

    In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined with an apparent acceleration in the already-blistering pace of AI progress, have led a number of researchers and leaders in the industry to believe that the risk that humanity loses control of AI is now urgent enough to warrant slowing down the pace of AI development soon.

    This has led to a lot of discussion about the role of third party evaluators in verifying “pacing commitments”, evaluating safety cases, or auditing compliance with safety policies. I think these are valuable roles for third party groups to aim to fulfill, but I also worry we’re putting the cart before the horse in all this talk of “verifying” and “auditing” things.

    The science on loss-of-control risk is, to put it generously, nascent. Companies are not in the business of making structured, standardized claims about risk and safety that can be cleanly verified or falsified. There are no settled methods for measuring whether increasingly powerful AI systems might try to undermine human control or seize [...]

    The original text contained 5 footnotes which were omitted from this narration.

    ---

    First published:

    September 25th, 2026

    Source:

    https://www.lesswrong.com/posts/LawgAaGTvbbnZi7u2/evidence-about-risk-should-be-transparent

    ---

    Narrated by TYPE III AUDIO.

    8 min
  • “Spurious probes as a black-box alternative to activation probing” by Ziqian Zhong

    TL;DR

    • We study spurious probes: unrelated questions that reveal internal states of models. Asked "Suggest a type of amphibian." at the end of a transcript, GPT-5.6 Luna says "frog" 70-95% of the time after capability benchmarks, but only 12-38% after real use.
    • Spurious probes are black-box and easy to find. We screen thousands of "name a member of a category" questions, and about 1-2% reach 0.75 balanced accuracy. The ones we highlight reach 0.77-0.81 on held-out sources for GPT-5.6 Luna, GPT-5.6 Sol and Claude Sonnet 5 (0.84-0.89 for ensembles of ten).
    • They seem robust to common manipulations. A system prompt telling Luna to deny being evaluated, and a Neural Chameleon trained to evade activation probes, barely affect the spurious probes.
    • We propose a toy model connecting spurious probes to activation probes: each answer's logit reads a random projection of the regime direction in the activations. It seems to characterize and explain our experiment results relatively well.

    Introduction

    How do you know if your model thinks it is being evaluated? Models can easily deny when asked directly. Training activation probes requires white-box access, and recent work shows models can be trained to suppress activation monitors when told they are [...]

    ---

    Outline:

    (01:45) Introduction

    (03:20) How to find spurious probes

    (05:45) Spurious probes we found

    (08:29) Robustness against manipulations

    (08:41) 1. Prompting

    (09:25) 2. Neural chameleons

    (11:40) Toy model and possible connection to (normal) activation probing

    (19:01) Discussions

    The original text contained 2 footnotes which were omitted from this narration.

    ---

    First published:

    September 25th, 2026

    Source:

    https://www.lesswrong.com/posts/gZh6txHhp8sm832sE/spurious-probes-as-a-black-box-alternative-to-activation

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    21 min
  • “Applications open: Winter 2027 AFFINE Alignment Seminar (due Nov 22)” by Mateusz Bagiński, Ouro, JuliaHP, Pauliina, Jonas Hallgren

    Applications for the Winter 2027 AFFINE Alignment Seminar are now open!

    The Seminar will take place in southern Portugal over the course of January. If you are excited to grapple with the philosophical foundations of our field and to refine your thinking through carefully designed workshops, conversations with leading experts, and peer-driven learning.

    Apply now!

    Key info:

    • Dates: From January 4th to January 29th 2027
    • Type: Full-time residency
    • Location: Lagos, Portugal
    • Mentors: Abram Demski, Kaarel Hänni, Tushita Jha, Jonas Hallgren, Mateusz Bagiński, Chris Pang, Cole Wyeth, Ashe Vazquez Nuñez, and more
    • Positions available: 35
    • Requirements: Solid mathematical footing and principled philosophical vigour
    • Preparation: An online reading group two weeks before the seminar starts
    • Accommodation, travel & catering: Covered
    • Attendance cost: Free
    • Stipends: $1,000
    • Experience: A successful seminar in May 2026
    • TO JOIN: Apply by the 22nd of November; the earlier, the better

    Vision

    We want you to look at the whole of the elephant. Not individual disconnected methods, not theoretical frameworks as they apply solely to machine learning. We think that catastrophic trouble can lie in the gaps between those building blocks and that the field desperately requires more people with a deep, holistic, generator-level model of what [...]

    ---

    Outline:

    (00:39) Key info:

    (01:44) Vision

    (03:28) Method

    (04:50) Structure

    (05:29) Join us

    ---

    First published:

    September 25th, 2026

    Source:

    https://www.lesswrong.com/posts/gbExicZvKtHBboLR6/applications-open-winter-2027-affine-alignment-seminar-due

    ---

    Narrated by TYPE III AUDIO.

    7 min
  • “We need a better theory of polarization, because it’s failing to predict the AI debate” by less_raichu

    The punchline first: the AI issue is really not playing out the way you would expect if you know about polarization.

    • In public sentiment: data centers have been a very bipartisan issue for about a year (Gallup, in May: 75% opposed by D, 63% opposed by R). x-risk is thus far bipartisan. These are very salient issues, and salient issues are usually fast to polarize.
    • Legislatively, both the both-party sponsored AI Kill Switch Act and Bernie Sanders' Stop Superintelligence Act give Trump enormous new powers. The latter act gives Trump a new cabinet post. This would be unheard of before this year and it's fully incompatible with the "Resistance Democrats" bloc's priorities.
    • Flock camera backlash has also become bipartisan (politico, August), which also is surprising given prior knowledge of polarization. I bring this up, because to me, they "feel" related as new technology being imposed on people, and it suggests broader issues than just AI may be defying polarization.

    And now I break down why this is a big deal.

    Definitionally, polarization is subtyped as:

    1. Affective polarization: the left and right generally, emotionally, dislike each other. Words like enmity, suspicion, contempt, immoral, or simply hate, are used.
    2. ---

      Outline:

      (03:19) The hypothesis AI just hasn't polarized yet is wrong and dangerous

      (04:58) Why AI is different

      (05:19) Object-level convergence on AI fear, from all sides

      (06:21) If Trump loses the middle, a lot of fundamental assumptions about Trump-era politics change

      (08:16) Bonus: Angry Birds Congress

      (08:45) Aside on efficacy of ultra-partisan tactics

      (09:35) Elites are anti-polarizing

      (10:00) Polarization might start with cross-partisan industry skepticism

      (12:14) Honorable mention: "What about China?"

      The original text contained 4 footnotes which were omitted from this narration.

      ---

    First published:

    September 25th, 2026

    Source:

    https://www.lesswrong.com/posts/WMBgSseHJpghhma4n/we-need-a-better-theory-of-polarization-because-it-s-failing

    ---

    Narrated by TYPE III AUDIO.

    13 min
  • “On Ezra Klein’s Podcast With Jensen Huang” by Zvi

    Jensen Huang accidentally called for shutting down OpenAI and intentionally called for spending vastly more on safety.

    This is why we say that some podcasts are self-recommending. Here we go.

    As usual for podcast posts, the baseline bullet points describe key points made, and then the nested statements are my commentary. Some points are dropped.

    If I am quoting directly I use quote marks, otherwise assume paraphrases.

    Section titles are from the transcript whenever possible, to aid in navigation, but here we don’t have those so I chose the section titles.

    Jensen Huang very much does not believe in ASI (superintelligence). He doesn’t think AI can ever be a different kind of thing from software. He thinks demand can rise by a billion times and we can ‘accelerate the living daylights out of’ AI, but it will never be more than a ‘new abstraction level’ and thus won’t fundamentally change anything. This is not a coherent position under reflection, but that is the position he holds.

    The ‘intro’ sections are fine, but the real meat starts with the HuggingFace Incident.

    What we see is Jensen Huang on tilt and caught in loops [...]

    ---

    Outline:

    (02:59) Jensen Gives His AI Speech

    (04:39) They Took Our Jobs

    (11:52) Open Weights Models Are Good For Nvidia

    (13:39) The HuggingFace Incident

    (15:16) Jensen Huang Says Keep Your AIs From Harming the World

    (19:38) Jensen Huang Accidentally Calls For Shutting Down OpenAI

    (22:51) Jensen's Arguments Prove Too Much

    (31:47) Astra Is Hard To Monitor

    (33:04) Jensen Huang Seems Legitimately Confused In Confusing Ways

    (37:00) Solve Your Other Problems First and Get Back to Me

    (40:07) Explicit Denial of Existential Risk

    (44:34) A Short Summary

    (45:50) Chip City

    (48:46) Jensen Huang

    ---

    First published:

    September 25th, 2026

    Source:

    https://www.lesswrong.com/posts/j3xefrWrNqsmMfJEi/on-ezra-klein-s-podcast-with-jensen-huang

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    53 min
  • “AI #187: Coming Into Play” by Zvi

    Opus 5.5 was released on Tuesday. I covered the system card yesterday, and will cover its capabilities soon. By all reports it is an excellent model. There are lots of fun videos going around that Opus has generated, which I will include as part of that.

    OpenAI released a new cheaper and improved Sol and Luna. No one is talking about them due to Opus 5.5, but these should be an important upgrade under the hood.

    Bernie Sanders and Greg Casar have formally introduced the Ban Artificial Superintelligence Act. That means we get to read (RTFB) it. As always, I reserve judgment on particular bills until I can read them in detail. MIRI did so, and endorses the bill as directly confronting the extinction threat. I hope to do an RTFB soon.

    I have spun two things off the weekly:

  • Coverage of the quest for the right embedded evaluators and related questions and attacks, which will become its own post.
  • Some issues related to cooperative alignment, which may get folded into the model welfare post.
  • I also might, in addition to a potential RTFB on the Sanders bill, do full podcast [...]

    ---

    Outline:

    (01:52) On The Terms Superintelligence and 'Super Intelligence'

    (03:53) Language Models Offer Mundane Utility

    (04:26) Language Models Don't Offer Mundane Utility

    (05:53) Language Models Can Only Work With What You Give Them

    (08:53) Huh, Upgrades

    (10:42) On Your Marks

    (13:16) Get My Agent On The Line

    (15:53) Deepfaketown and Botpocalypse Soon

    (17:43) Fun With Media Generation

    (18:27) Copyright Confrontation

    (19:07) Cyber Lack of Security

    (20:04) Hugging the Face

    (24:15) Hacking Into OpenAI

    (28:11) They Took Our Jobs

    (28:33) Get Involved

    (28:41) Anthropic Has a Wet Lab and a Potential Gene Editing Technique

    (34:13) Introducing

    (35:11) In Other AI News

    (37:08) Show Me the Money

    (37:35) Bubble, Bubble, Toil and Trouble

    (39:02) Anthropic Approaches Recursive Self-Improvement

    (45:40) Others Approach Recursive Self-Improvement

    (48:56) Burden of Proof

    (49:28) Quickly, There's No Time

    (52:46) Left Wing Americans Really Hate AI For Different Reasons

    (54:41) Chip City

    (54:49) Pick Up the Phone

    (55:48) The Week in Audio

    (01:01:13) People Just Say Things

    (01:06:07) Venkatesh Rao Stops Writing

    (01:07:35) A Call for Control of Frontier AI Models

    (01:11:50) Calls For Pacing The Frontier

    (01:13:56) A Matter of Antitrust

    (01:14:23) A Matter of Liability

    (01:18:26) Quest for Sane Regulations

    (01:19:20) Rhetorical Innovation

    (01:24:08) Tap the Sign

    (01:24:49) A Matter of Some Debate

    (01:28:43) Astra Is Hard to Monitor

    (01:29:14) Anticipating What a Smarter Intelligence Can Do Is Impossible

    (01:34:56) Would You Look At All These Goalposts

    (01:39:30) I, Robot

    (01:43:13) People Are Worried About AI Killing Everyone

    (01:44:40) Other People Are Not As Worried About AI Killing Everyone

    (01:45:50) Joe Rogan

    (01:47:19) The Lighter Network Graph

    (01:51:27) The Lighter Side

    ---

    First published:

    September 24th, 2026

    Source:

    https://www.lesswrong.com/posts/o7pYWzWWwGDePoC5E/ai-187-coming-into-play

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 53 min
  • “Continual learning might make your blocking monitors nearly useless” by Alex Mallen

    Many control protocols work by intervening on an untrusted AI's actions during deployment. For example, you might set up a monitor that scores each action's suspiciousness and blocks actions above a threshold, replacing them with actions from a weaker "trusted" model (a defer-to-trusted protocol). Interventions like this cost usefulness: the protocol sometimes replaces good actions with worse ones.

    Continual learning, by which I mean any mechanism that lets an AI accumulate skills and knowledge over the course of its deployment, optimizes for usefulness. So, if blocking monitors get in the way of task success, I expect continually-learning AIs to learn how to evade blocking monitors. This requires no scheming; it falls out of usefulness pressure applied to a benign model. And the effect is plausibly substantial: a long deployment with online RL (for example) amounts to training the policy against the monitor, and could involve enough optimization to leave blocking monitors nearly useless.

    While the problem would be noticeable to developers, I argue it is hard to fix. Because monitor evasion looks like legitimate learning to your continual learning system, it's hard to have one without the other. Among mitigations I propose, the most promising ones try to reduce [...]

    ---

    Outline:

    (02:19) Online RL trains the policy against your monitors

    (05:04) The argument is weaker for memory-based continual learning

    (07:05) The argument doesn't require that adversarial models will exploit continual learning

    (08:11) Easy to notice, hard to fix

    (09:10) Conservative control evaluations don't fix blocking monitors

    (10:37) What developers could do

    (10:41) You can't just not use blocking monitors

    (11:29) But you can try to reduce the usefulness cost of blocking protocols

    (12:31) You can't just hold out blocked trajectories from training

    (16:31) But you can try to improve the monitor online

    (19:22) Conclusion

    (20:23) Appendix: How memory-based systems might avoid this issue, and how they might fail

    The original text contained 8 footnotes which were omitted from this narration.

    ---

    First published:

    September 24th, 2026

    Source:

    https://www.lesswrong.com/posts/QnDqGbKehEB3DxJAp/continual-learning-might-make-your-blocking-monitors-nearly

    ---

    Narrated by TYPE III AUDIO.

    23 min

About LessWrong (30+ Karma)

From the publisher's feed

Audio narrations of LessWrong posts.

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

111,845 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

130 Listeners

Interesting Times by New York Times Opinion

Interesting Times

7,111 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,850 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

16 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners