LessWrong (30+ Karma)

LessWrong (30+ Karma)

Download on the App Store

LessWrong (30+ Karma) episodes

  • “The Anatomy of a Chinese AI Researcher” by CMLKevin

    The Chinese AI researcher has read the Three Body Problem series of sci-fi novels since high school, and understand the concept of existential risk vaguely.

    He is fascinated by Ye Wenjie, the researcher that turned against humanity in that book, and decides that in the future if AI progress leads to a superior intelligence, he might be tempted to become Ye if there's no good alternative.

    He performs the duties of capabilities research in a Chinese frontier lab, seeking to one day achieve parity with Western companies, though he knows this is difficult. He has a mentality of hillclimbing, believing that the progress of a future technology is highly uncertain and even unknowable, and so him and his peers could only tread one step at a time.

    He looks at the western world and sees what is typical when a great technology is developed: the first mover will decide to impose restrictions to further their lead, while latecomers should use whatever means necessary to widen access to the whole world. He thinks of the AI chip restrictions as evidence of this.

    He uses Anthropic and OpenAI models regularly in his day to day work. He [...]





    ---

    First published:

    September 19th, 2026

    Source:

    https://www.lesswrong.com/posts/qmxkHm2dTLKG6GZ6i/the-anatomy-of-a-chinese-ai-researcher

    ---

    Narrated by TYPE III AUDIO.

    4 min
  • “Why I Stay Off Twitter” by jefftk

    I avoid Twitter (𝕏) for similar reasons to drugs: I think it
    would change me for the worse, and I would be unable to give it up.

    After staying off Twitter reasonably successfully for years, I
    cross-posted my AI
    Tweets there a few weeks ago. I had something very Twitter-shaped
    to say, and I thought it was important to get out, so I do
    think this was worth it. And it all went well: none of this is
    complaining about the comments I got there.

    Coming back a few times to check notifications, however, it's been
    very good at baiting me: Tweets that are confidently wrong in cases
    where I have relevant and uncommon knowledge. The pull to dive in and
    share what I know is very strong! Then this bleeds over to the far
    broader case where people are wrong, and you have a large potential
    time sink.

    If it were just the time sink, I'd stop resisting. I spend some
    time on HN and Reddit, and to the extent that Twitter could substitute
    for that by showing me things I was more interested in, that wouldn't
    be an issue. The real problem [...]

    ---

    First published:

    September 19th, 2026

    Source:

    https://www.lesswrong.com/posts/tvwtwgcujTfep4HgY/why-i-stay-off-twitter

    ---

    Narrated by TYPE III AUDIO.

    4 min
  • “NYT Editorial Board Comes Out Against Extinction” by Ben Pace

    (Archive link)

    The NYT editorial board's article on AI (archive link) is far better than I'd expected, but at the same time not all I'd hoped for.

    The title sets off very well: "Humanity Has Avoided Apocalypse Before. Let's Do It Again." It is truly excellent to see the extinction threat from loss of control be mainlined.

    A quick gloss of their policy requests: an AI Commission in government, licensing requirements for AI companies, an AI "constitution" written by the US Government incorporated into AIs, mandatory watermarks/identifiers on all AI content, mandatory independent testing for AI models before release, and a government agency to investigate accidents. Internationally, they call for tightening export controls, limiting China's access to semiconductors, and ultimately negotiating an international slowdown with China and an international framework for AI oversight.

    These are all steps in the right direction—of taking AI seriously. That said, it isn't clear if the licensing is required for training or for selling AIs. The idea that constitutional AI "would ensure alignment with human values" is of course not remotely true. And mandatory testing should apply to all models trained, not all models released, of course, and this is a glaring oversight. But [...]

    ---

    First published:

    September 19th, 2026

    Source:

    https://www.lesswrong.com/posts/gDQzntJCusNbshWyD/nyt-editorial-board-comes-out-against-extinction

    ---

    Narrated by TYPE III AUDIO.

    4 min
  • “Common mistakes in AI safety group organizing” by Nikola Jurkovic

    Back in the day, I was a very active AI safety group organizer. I commonly notice people making the same mistakes across many clubs. I have written down a list of some of these mistakes hoping people will avoid them in the future:

    • Reading groups often require that people read things before meetings. This is a mistake. People often don't do the readings. And the lack of common knowledge that everyone has read the reading degrades the conversation quality.
      • Instead, have longer meetings, serve food (so, lunch/dinner meeting slots), and read during the actual meeting.
    • Reading groups often don't sort people into cohorts properly. Mainly, they fail at clustering people into clusters of roughly equal ML knowledge and age. Grad students don't want to discuss a paper with freshmen. People with lots of ML knowledge don't want to discuss a paper with people with no ML knowledge.
      • Instead, group people with people similar to them in ML knowledge and age.
    • Reading groups often rely on digital materials instead of physical printouts. Screens are distracting and there is no common knowledge that people are paying attention.
      • Neatly print every reading ahead of time instead.
    • Clubs [...]

    ---

    First published:

    September 19th, 2026

    Source:

    https://www.lesswrong.com/posts/XFzqDJjAJBt8fkn8i/common-mistakes-in-ai-safety-group-organizing

    ---

    Narrated by TYPE III AUDIO.

    4 min
  • “The AI Risk Network” by derelict5432

    Most conversations about AI risks seem like people are talking past each other. There's a lot of strawmanning. This is largely because the issue is complex, and in order to make it manageable, the concepts get oversimplified. The most common example of this is p(doom), collapsing extreme outcomes into a single variable and focusing on that. Critics happily jump on the fact that there are many other possible bad outcomes, or that 100% extinction rates are an extremely high bar. There's some legitimacy to this.

    So here I want to try to grapple with the the complexity of the larger web of issues surrounding bad AI outcomes by presenting The AI Risk Network. If you’re interested in this topic, bear with me. It might be a bit of a slog.

    First I want to contrast this approach with others. Liron Shapira has what he calls The Doom Train, a linear progression through various dependencies or thresholds that eventually lead to human extinction, with various ‘stops’ along the way where the skeptic can get off.

    Shapira uses this as a discussion guide to focus on particular points where the skeptic gets off the train and exits belief in the extreme [...]

    ---

    First published:

    September 19th, 2026

    Source:

    https://www.lesswrong.com/posts/XundBXqKSo3bB2A6a/the-ai-risk-network

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    10 min
  • “Anthropic Looks At Some Of Its Alignment Problems” by Zvi

    Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI.

    There will also be a METR investigation of these incidents, which unlike the investigation done at OpenAI will be untimed.

    Table of Contents

  • Our Two Problems.
  • First the Good News.
  • We’d Just Like To Ask You a Few Questions.
  • Internal Research Model On The Fence.
  • Opus 4.7.
  • Opus 4.6 Checkpoint.
  • Holy **** That Thing's Real?
  • I Thought I Saw a Pussycat.
  • If This Was Real You Would Never Tell Me It Was Real.
  • New Eval Who Dis.
  • Hacker Opus.
  • Monitoring the Situation.
  • Overcoming Bias.
  • The Anthropic Alignment Problem.
  • Paths Forward.
  • Our Two Problems

    Anthropic: Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents:

  • biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet
  • recklessness, or a willingness to take harmful actions in the narrow pursuit [...]
  • ---

    Outline:

    (00:33) Our Two Problems

    (02:26) First the Good News

    (03:02) We'd Just Like To Ask You a Few Questions

    (04:12) Internal Research Model On The Fence

    (07:28) Opus 4.7

    (08:12) Opus 4.6 Checkpoint

    (09:49) Holy **** That Thing's Real?

    (11:45) I Thought I Saw a Pussycat

    (19:28) If This Was Real You Would Never Tell Me It Was Real

    (21:19) New Eval Who Dis

    (26:32) Hacker Opus

    (30:15) Monitoring the Situation

    (31:38) Overcoming Bias

    (33:40) The Anthropic Alignment Problem

    (35:53) Paths Forward

    ---

    First published:

    September 19th, 2026

    Source:

    https://www.lesswrong.com/posts/ggFx5Wb3Hi4pJsueK/anthropic-looks-at-some-of-its-alignment-problems

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    38 min
  • “You Should Apply to Inkhaven” by Tomás B.

    Inkhaven is a writers residency in Berkeley, in which the only requirement is you have to publish 500 words each and every day. Though I always had some confidence in my ability to write, I never actually did it much until I applied to Inkhaven. I had finished only two short stories before I applied: The Maker of MIND and The Liar and the Scold. And it was them I used in my application.

    In the roughly twelve months since I was accepted, I have written thirteen, and even some half-finished things that will never see the light of day. And this isn’t including the essays and micro-fiction I wrote during the fellowship. By the metric of getting me to write more, Inkhaven was a great success. And would have been worth it even if I had a miserable time.

    Despite a slight proclivity for having miserable times, I found myself unable to do so for long at Inkhaven. I rarely write utopias, and when I do they curdle by the time the story ends. But I suspect utopia will feel a lot like Inkhaven did for me once I got settled. You would think putting a bunch [...]

    ---

    First published:

    September 18th, 2026

    Source:

    https://www.lesswrong.com/posts/CKkB9MqsBgAobFtPS/you-should-apply-to-inkhaven

    ---

    Narrated by TYPE III AUDIO.

    4 min
  • “Pretraining data, not verifiability, is why LLMs are especially good at math (and coding)” by Steven Byrnes

    Follow-up to: “LLMs are (still) mostly powered by imitative learning, not RL”

    A common take I’ve been hearing is: “LLMs are especially good at math because math is easy to verify”. But that story doesn’t make much sense.

    • For one thing, “easy to verify” only matters for the RL part of LLM training pipelines, and the leading LLM companies have said that they spend very little effort on RL-for-math.
    • Worse, to the extent that the companies are doing RL-for-math, it's RLAIF, not RLVR. So really, the phrase “math is easy to verify” amounts to “LLMs are very good at judging math arguments”. But that's begging the question! Why are pretrained LLMs so much better at judging math arguments than judging, say, fiction writing? We still need an answer.

    So here's a different theory, in the framework of my earlier post “LLMs are (still) mostly powered by imitative learning, not RL”:

    LLMs are especially good at math because almost everything in the math literature is correct. Read a random sentence in a random math paper in the research math literature, and you can be >99% confident that the sentence is true. So if LLMs do what they do best—imitative [...]

    The original text contained 3 footnotes which were omitted from this narration.

    ---

    First published:

    September 18th, 2026

    Source:

    https://www.lesswrong.com/posts/xvdngZAqFZfek7KGH/pretraining-data-not-verifiability-is-why-llms-are

    ---

    Narrated by TYPE III AUDIO.

    5 min
  • “You don’t need a union to go on strike” by sudo-nym

    I'm mostly hoping this somehow gets sent to a privately disgruntled frontier lab employee, but it would also be cool to expand other people's minds on the way there.

    I read through Ethical AI Departures and would like to note that only a few of them have gotten extensive media coverage and none of them have actually effectively gotten the frontier labs to stop, and that collectively signed letters by employees have historically not done much either.

    I read Dear God, Please Do Not Resign In Protest and wanted to point out that leftists have a mature and relatively reliable set of strategies to address the problem of how to get a lot of people to stop working in protest at the same time.

    Then I did a search of LW to see if someone else brought unions up already, read What if AI safety labs unionized?, and flinched at the repeated citation of legal reasons why a union isn't the correct legal structure. So no, what you want right now isn't an official, bureaucratic union. In fact, that would probably slow things down too much.

    But I've done enough work with union people to know that you don't [...]

    ---

    First published:

    September 17th, 2026

    Source:

    https://www.lesswrong.com/posts/erd4NSztYMynbuTKw/you-don-t-need-a-union-to-go-on-strike

    ---

    Narrated by TYPE III AUDIO.

    3 min
  • “Stopgap Measures to Address Immediate AI Security Threats” by Andrea_Miotti, Gabriel Alfour

    Most people know AI as the technology behind chatbots like ChatGPT. However, what the top AI companies are explicitly aiming for is something else entirely: superintelligent AI. That means AI that can fully replace and outmatch humans at any task, including in domains like hacking, social engineering, and military operations. Such an AI system, if developed, could autonomously overpower any country's national security forces.

    No company, no government, no individual knows how to keep such a system under human control. This is why the world's leading AI experts, Nobel Prize winners, and even the CEOs of the top AI companies warn that the development of superintelligence threatens humanity with extinction, and why more than 800 scientists, former military leaders, and public figures have called for a prohibition on developing superintelligence.

    This is not a distant prospect: AI companies such as OpenAI and Anthropic are investing billions of dollars into superintelligence and aiming to develop it within the next few years. Former Anthropic and OpenAI researcher Jacob Coxon, who resigned last week, stated that AI companies are “racing straight to self-improving superintelligence and gambling with our lives” and that people at the companies themselves believe it “could kill [...]

    ---

    Outline:

    (04:34) Secure Weapons-Grade AI Against Theft by Adversaries

    (07:46) Necessary Measure: Registration

    (08:43) Sufficient Measure: Government Security Testing

    (09:42) Thorough Measure: Development Requires Government Authorization

    (10:52) Criminal Liability for Leaks During AI Gain-of-Function Research

    (14:09) Necessary Measure: Team Liability

    (14:46) Sufficient Measure: Chain of Command Liability

    (15:21) Thorough Measure: Company Liability

    (16:00) Kill-Switches to Contain Critical AI Incidents

    (19:08) Necessary Measure: Company Kill-Switch

    (19:58) Sufficient Measure: Infrastructure Kill-Switch

    (20:53) Thorough Measure: International Kill-Switches

    (22:24) Conclusion

    ---

    First published:

    September 18th, 2026

    Source:

    https://www.lesswrong.com/posts/LqBAxFdyAiybnPL8e/stopgap-measures-to-address-immediate-ai-security-threats

    ---

    Narrated by TYPE III AUDIO.

    25 min

About LessWrong (30+ Karma)

From the publisher's feed

Audio narrations of LessWrong posts.

More shows like LessWrong (30+ Karma)

The Daily by The New York Times

The Daily

111,845 Listeners

Astral Codex Ten Podcast by Jeremiah

Astral Codex Ten Podcast

130 Listeners

Interesting Times by New York Times Opinion

Interesting Times

7,111 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,850 Listeners

AI Article Readings by Readings of great articles in AI voices

AI Article Readings

4 Listeners

Doom Debates! by Liron Shapira

Doom Debates!

16 Listeners

LessWrong posts by zvi by zvi

LessWrong posts by zvi

2 Listeners