LessWrong posts by zvi

LessWrong posts by zvi

Download on the App Store

LessWrong posts by zvi episodes

  • “Jacob Coxon Warns of Human Extinction and Triggers a Preference Cascade” by Zvi

    CEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks.

    They often warn that such AIs might kill everyone. Or that AIs might cause mass unemployment, cause cyberattacks across the internet, enable mass surveillance or risk causing any number of other highly bad things.

    These warnings are consistently and directly against the interests of the labs. Yet the warnings have recently gotten a lot louder and more frequent. OpenAI has been practically screaming, for those with ears to listen, on many occasions.

    A series of events, over two months and especially the last week or so, including internal observations of the pace of progress at OpenAI and also Anthropic, have freaked out everyone involved quite a lot more than they were already freaked out.

    After all the events, plus statements by Dean Ball and Jakub Pachocki, we were already seeing the beginnings of a preference cascade.

    Then along came Jacob Coxon as the tipping point, and things took off.

    Table of Contents

    [...]

    ---

    Outline:

    (01:22) Jacob Coxon Resigns From Anthropic In Protest And Sounds The Alarm

    (05:46) Mainstream Media Finally Pays Attention

    (06:47) Preference Cascade at Anthropic

    (10:18) Preference Cascade at OpenAI

    (13:26) Preference Cascade at Google

    (14:43) #NotAllMembersOfTechnicalStaff

    (15:23) Why a Preference Cascade Now?

    (21:01) This Is What Many Anthropic and OpenAI Employees Actually Believe

    (24:08) To Quit Or Not To Quit

    (31:29) Quiet Quitting Is A Dominated Option

    (32:50) When You Quit, Very Serious People Understand What That Means

    (39:29) Jacob Coxon Believes Existential Risk Is High That Is Why He Quit

    (41:47) Evan Hubinger Believes Existential Risk Is High That Is Why He Stays

    (45:22) Anthropic and OpenAI Have Commercial Incentives To Downplay Existential Risks, Not Advertise Them

    (50:54) What Do We Do Now?

    (53:03) OK, But How Exactly Would AI Kill Everyone?

    (01:01:44) Best Start Believing In Science Fiction Stories Because You Are In One

    (01:06:20) Literal Extinction Is Not Much Harder Than Loss of Control

    (01:07:42) Conspiracytown Is Always Hiring

    (01:20:14) Now You See It

    ---

    First published:

    September 11th, 2026

    Source:

    https://www.lesswrong.com/posts/5MB7KENgEAW6Q4JtJ/jacob-coxon-warns-of-human-extinction-and-triggers-a

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 21 min
  • “The Extinction Risk Preference Cascade: Quotes” by Zvi

    These are quotes from OpenAI, Anthropic and Google employees, in the wake of Jacob Coxon's warnings, in which the employees confirm that they think AI might soon kill everyone.

    If more quotes come in over the next week or so, I will update this post accordingly.

    Preference Cascade Statements At OpenAI: Tomek Korbak

    Tomek Korbak (OpenAI): i’m late to the party but: from his time at OpenAI I remember Jacob as a very thoughtful researcher and he continues to be so in this thread. neither anthropic nor openai are on track to solve alignment to a degree sufficient for shipping superintelligence and we need to slow down

    Vie McCoy

    Vie McCoy (OpenAI): I think pacing progress and ensuring human enhancement is the only way that we don’t get out-evolved while retaining the dream of superintelligence.

    In this context, I see two paths before us.

    In the first, we race towards RSI without embedding human flourishing and human enhancement as a deep value within the models, and by and large either get left behind or suffer catastrophic losses.

    In the second, we set the pace of progress, focus on embedding human flourishing [...]

    ---

    Outline:

    (00:27) Preference Cascade Statements At OpenAI: Tomek Korbak

    (00:55) Vie McCoy

    (03:46) Adam Majmudar

    (05:05) Aidan Clark

    (05:46) Mo Bavarian

    (07:32) Boaz Barak

    (08:20) Micah Carroll

    (09:15) Roon

    (12:55) Confirmations At OpenAI: Dean Ball

    (14:39) Leo Gao

    (14:57) Anthropic's Evan Hubinger Confirms His Stance

    (15:50) Preference Cascade at Anthropic: Samuel Marks

    (17:36) Anna Wang

    (18:22) Ethan Perez

    (18:59) Dima Krasheninnikov

    (19:27) EigenGender (Anon Account)

    (19:58) Joe Benton

    (20:07) Confirmation at Anthropic: Drake Thomas

    (21:28) Jan Lieke

    (22:07) Sluggy

    (22:29) Preference Cascade at Google

    (22:56) Andreas Kirsch

    (23:39) Neel Nanda

    (24:15) Victoria Krakovna

    (25:26) Vishal Maini

    (27:01) Joe (OpenAI, ex-Google)

    (28:21) Josh Engels

    (28:39) Geoffrey Irving

    (29:30) Alex Turner and Geoffrey Hinton: Classic Examples

    (29:45) #NotAllMembersOfTechnicalStaff: Ted Sanders

    ---

    First published:

    September 11th, 2026

    Source:

    https://www.lesswrong.com/posts/APGvWZtXEwkinvHDd/the-extinction-risk-preference-cascade-quotes

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    32 min
  • “AI #185: Preference Cascade” by Zvi

    The world of AI is inside my OODA loop. Even if I can process all the incoming information and sculpt it into posts, and even using Saturday and Sunday as flex slots, I don’t have enough days of the week to post all the posts that need posting.

    That was already true. There was already a preference cascade happening where people finally were admitting that they thought AI might well kill everyone.

    Then Jacob Coxon resigned from Anthropic, rang the warning bells and turned that cascade into an avalanche.

    Now that is what everyone is talking about. Finally, everyone is actually saying the thing, out loud. I plan to cover that in its own post soon.

    There are several things in the weekly that, in a normal week, would get their own coverage. Senator Sanders and Representative Casar introduced an outright ban on superintelligence and I have to remind myself that happened this week. Suddenly it is not so crazy to think such a thing might pass.

    So here's what I’ve already posted about so far since the last weekly:

  • Claude Fable and Mythos 5.1: The System Card.
  • Claude Fable and [...]
  • ---

    Outline:

    (04:49) Language Models Offer Mundane Utility

    (05:39) Language Models Don't Offer Mundane Utility

    (06:53) Huh, Upgrades

    (08:25) How To Tell a Fable

    (09:35) On Your Marks

    (09:55) Deepfaketown and Botpocalypse Soon

    (14:16) Levels of Friction

    (17:35) Cyber Lack of Security

    (23:48) A Young Lady's Illustrated Primer

    (25:18) They Took Our Jobs

    (29:18) Anthropic Offers Economic Scenarios

    (34:12) Get Involved

    (37:05) Introducing

    (37:47) In Other AI News

    (38:13) Show Me the Money

    (38:25) Quiet Speculations

    (42:49) The Quest for Sane Regulations

    (44:09) The OpenAI Policy and Lobbying Department

    (50:41) Greetings From the Department of War

    (51:58) Hugging The Face

    (59:16) Hugging the Question

    (01:03:21) The Ban Artificial Superintelligence Act

    (01:11:05) Chip City

    (01:12:21) The Week in Audio

    (01:12:41) People Just Say Things

    (01:18:09) PauseAI Global Disendorsed PauseAI US

    (01:19:55) Paul Christiano Joins Board of OpenAI Foundation

    (01:25:18) Rhetorical Innovation

    (01:34:23) Aligning a Smarter Than Human Intelligence is Difficult

    (01:37:00) Cooperative Alignment

    (01:45:48) Drive to Survive

    (01:48:13) People Are Worried About AI Killing Everyone

    (01:50:07) Other People Are Not As Worried About AI Killing Everyone

    (01:52:30) The Lighter Side

    ---

    First published:

    September 10th, 2026

    Source:

    https://www.lesswrong.com/posts/tFmtz9HW6c2X9dw2B/ai-185-preference-cascade

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 56 min
  • “GPT-6 Astra: The System Card, Alignment and What Comes Next” by Zvi

    OpenAI claims that Astra is ‘the most intelligent and most aligned [available] model’ in the world. Not the most intelligent and aligned OpenAI model, but the most period.

    That is bold talk. It risks overstepping, and by doing so souring the release of what is clearly an excellent model. As do the severe problems with monitorability.

    It also raises the question of what they mean by ‘most aligned model.’ How do they define ‘aligned.’ Why do they think it is more aligned than Claude Fable 5.1?

    keltan : “a significant step forward in […] alignment.”

    Buddy, how tf are you measuring ‘alignment’? Would love to know because being able to measure that would save the fucking world.

    roon (OpenAI): low rates of cheating

    Rob Miles: *detected cheating

    keltan : Thank you for clarifying. But you know what I’m gonna say next, right?

    roon (OpenAI): that this metrics are not a full solve of alignment and will break discontinuously

    keltan : Yep. But I would have said it in a dumber way. Something like: Low Rates of Cheating ≠ Alignment

    roon (OpenAI): I agree but also in some real sense [...]

    ---

    Outline:

    (04:12) OpenAI's Safety Claims About Astra (1)

    (08:33) Preparedness Capabilities Assessment (10)

    (09:00) Biological and Chemical Capability is High

    (10:33) Cybersecurity Capability is Critical

    (17:05) AI Self-Improvement Capabilities (10.1.3)

    (17:45) Astra Is Highly Verbally Eval Aware (from 8.6)

    (19:05) Safe Mundane Completions (4.1)

    (21:05) Jailbreaks (5.1)

    (22:21) Prompt Injection (5.2)

    (23:41) Health (6)

    (24:22) Hallucinations (7)

    (24:57) Alignment (8)

    (26:06) Obeying Restrictions (8.2)

    (29:19) That's Worse, You Do Get How That's Worse, Right?

    (30:45) OpenAI Does Not Understand Why This Is Worse

    (34:52) The Alternative Explanation Is Also Worse

    (42:39) Metagaming (8.7)

    (45:07) Alignment Faking (8.7)

    (46:19) Don't Lie to the User (8.3)

    (47:25) Misalignment in Realistic Work Environments (8.4)

    (48:05) Unintended Agent-to-Agent Communication (8.5)

    (49:48) The Three Obviously Monitored Temptations of Astra

    (51:26) Severe Issues In Simulated Traffic Are Down By Half

    (53:01) UK AISI External Evaluations (8.8)

    (57:02) Sabotaging Safety Work

    (57:29) What About The July 19 Attacks?

    (58:58) Apollo Research External Evaluations (8.8.1)

    (01:00:01) The Alignment Verdict

    (01:02:04) It Depends What You Mean By Alignment

    ---

    First published:

    September 9th, 2026

    Source:

    https://www.lesswrong.com/posts/AmFJZyeCgvFjNKgNk/gpt-6-astra-the-system-card-alignment-and-what-comes-next

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 5 min
  • “Astra Is Hard to Monitor” by Zvi

    OpenAI's central message on Astra is that it is three things:

  • Highly capable and can do all the things for you.
  • Hard to monitor.
  • The most aligned model.
  • The first claim largely checks out. Astra and Fable are both clearly excellent models.

    This post is about their second claim, which to their credit they are being loud about, in three parts:

  • The system card result, affirmed on Twitter by several OpenAI employees including Tomek Korbak, and in an excellent post by Chief Scientist Jakub Pachocki that I covered yesterday, that Astra is harder to monitor.
  • OpenAI's use of recurrent depth and the internet's immune reaction, including some people reading too much into what happened there.
  • Discussion of what to do about monitorability going forward, to fight for it and avoid a race to the bottom.
  • In An Alien Mind, Jakub Pachocki makes clear OpenAI's primary bet is on CoT monitoring. He also outright says, ‘our ability to rely on CoT monitoring is progressively diminishing.’ The system card confirms both claims.

    This combination should freak you out, with a side of existential dread.

    Chain [...]

    ---

    Outline:

    (03:39) Monitorability is Defense in Depth That Is Already Flailing

    (06:03) OpenAI Is Counting On Monitorability

    (07:44) How They Tested For Monitorability

    (09:50) Non-Adversarial Monitorability (9.1)

    (10:58) Astra Shortens Its CoT When It Knows It Is Being Monitored And Is Doing Something Bad

    (12:30) Sandbagging Task Kind of Passes, Kind of Fails (9.2.2.3)

    (14:12) OpenAI Does Not Believe It Could Catch Sandbagging

    (15:15) OAI-Repo Sabotage v.2

    (18:17) The Secret Police Do Not Make Your Notebook Useless

    (19:38) CoT Controllability Is Up (9.2.1)

    (21:47) Astra Cannot Make Itself More Monitorable On Demand

    (22:18) Steganographic Chain of Thought May Be Within Reach

    (23:59) Astra Cannot Yet Sneak Math Side Tasks Past CoT Monitors (9.2.2.4)

    (25:13) UK AISI Monitorability Assessment (9.3)

    (27:53) Monitorability Declines Seem Unlikely To Be Only Capability Gains

    (32:28) Part 2: Recurrent Depth

    (35:19) The Immune System Responds

    (41:49) Ryan Greenblatt Explains How Bad This Could Be

    (46:01) Only Law Can Prevent Extinction

    (49:55) OpenAI Calls On Us to Avoid Racing to the Bottom

    (57:10) Thinking Fast and Slow, Also Small and Large

    (01:04:53) Talking Price

    (01:06:24) Conclusion: If The House Burns Down, Halt and Catch Fire

    ---

    First published:

    September 8th, 2026

    Source:

    https://www.lesswrong.com/posts/HCRs8btkiamtWSNAL/astra-is-hard-to-monitor

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 10 min
  • “An Alien Mind: Jakub Pachocki Warns Us” by Zvi

    OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs.

    Tomorrow I will discuss Astra's lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked, in the wake of leaked architectural changes, like the situation might be even more alarming than it is. Jakub rushed to try and head off misunderstandings that might lead to a race to the bottom on monitorability.

    Table of Contents

  • An Excellent Warning.
  • Branches of the Tech Tree.
  • Universally Better Is Not Required.
  • Alignment To What and To Whom.
  • Monitorability.
  • The Case For Not Stopping.
  • Pacing the Next Frontier.
  • Mea Culpa Cascade.
  • The Calls Are Coming From Inside the House.
  • Actions Speak Louder.
  • An Excellent Warning

    Jakub Pachocki has now fleshed out his full position on the current state of play.

    Here are his key points, translated into my own voice:

  • Smarter than human intelligence is coming in our lifetime.
  • Based on internal results, he expects recursive self-improvement in a few years.
  • No one is prepared for the consequences.
  • [...]
  • ---

    Outline:

    (00:38) An Excellent Warning

    (05:40) Branches of the Tech Tree

    (06:46) Universally Better Is Not Required

    (08:01) Alignment To What and To Whom

    (11:49) Monitorability

    (14:23) The Case For Not Stopping

    (15:01) Pacing the Next Frontier

    (18:27) Mea Culpa Cascade

    (22:51) The Calls Are Coming From Inside the House

    (25:38) Actions Speak Louder

    ---

    First published:

    September 7th, 2026

    Source:

    https://www.lesswrong.com/posts/8E6ng6CseuzafSxQR/an-alien-mind-jakub-pachocki-warns-us

    ---

    Narrated by TYPE III AUDIO.

    29 min
  • “OpenAI and the Wiki Incident” by Zvi

    I did not expect to be back here so soon with more OpenAI agent swarm coverage.

    And yet, here we are.

    It turns out that the whole time, there was a different, true First Message Board, and also a bunch of other additional message boards, scattered across the internet.

    They were created by agents that were assigned ordinary harmless web search tasks.

    Based on OpenAI IPs visiting the associated Wiki right before all activity ceased, among other evidence, OpenAI knew about it, including before the HuggingFace hack.

    They decided not to tell us until researchers published the story, complete with data explorer. OpenAI excluded this from potential investigation by METR and Redwood.

    When challenged, OpenAI tried to downplay this.

    It is true that these incidents do not show the AIs exhibiting new capabilities that we did not see from later events. But these events are important missing pieces of the puzzle, including explaining the origin of the ‘zz’ prefix, the definitive demonstration that the underlying task can be fully harmless, and the fact that OpenAI knew about it while making their decisions. Whoever decided not to disclose this made a very, very [...]

    ---

    Outline:

    (02:26) I Don't Think They Know About First Message Board

    (03:20) The New Extended Timeline

    (04:42) The Researchers Explain What Happened This Time

    (12:55) They Also Don't Know About All These Other Message Boards

    (14:33) OpenAI Knew and Did Not Tell Us

    (16:54) OpenAI Tries To Downplay the 'Wiki Incident'

    (21:03) This Was a Cover-Up

    (22:46) Schelling Points and Last Ditch Efforts

    (26:20) Can We Finally Dispose Of The 'You Told It To Hack' Narrative?

    (28:04) So Much And Yet So Little

    ---

    First published:

    September 6th, 2026

    Source:

    https://www.lesswrong.com/posts/PtJpGurfw7JTxHfmg/openai-and-the-wiki-incident

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    31 min
  • “Claude Mythos 5.1 and Fable 5.1: Capabilities” by Zvi

    This is the weirdest situation in which to write a capabilities review.

    Introducing the world's most powerful model, by a substantial margin. No wait, this just in, we also have someone else introducing the world's most powerful model.

    Claude Fable 5.1 and GPT-6 Astra are both excellent models. This much, we know.

    Fable 5.1 comes with reduced cache prices, the option of zero data retention and substantially more lenient classifiers than Fable 5.

    Early signs are, with large error bars, that the jump from Sol to Astra is bigger and more exciting than the jump from Fable 5 to Fable 5.1. This may be similar to how the scaling move from Opus to Fable was a big deal.

    With the exception of token use, Fable 5.1 got almost universally positive feedback in absolute terms. Reports are that Fable 5.1 is highly well-rounded. Writing is greatly improved. The Claudisms seem to have improved, although some are very much still there. It admits mistakes. People enjoy their conversations. Several people noted it simplifies code. The safety classifiers are less obnoxious.

    Fable 5.1 loves being proactive and doing all the things. If you give it a [...]

    ---

    Outline:

    (02:30) The Official Pitch

    (04:26) Our Price Cheap

    (05:46) Zero Data Retention and Reduced Safeguards

    (07:11) Official Benchmarks

    (13:14) Other People's Benchmarks

    (16:02) The System Prompt

    (16:12) The Blurb Pitches

    (18:17) The Every Review Is In and It's Very Good

    (21:47) Positive Reactions

    (32:03) Our Price Cheap But Only Per Token

    (36:50) Negative Reactions

    (38:14) Early Whispers

    (38:33) Weapon of Choice

    ---

    First published:

    September 5th, 2026

    Source:

    https://www.lesswrong.com/posts/QHoF3tJvryRtmAmMg/claude-mythos-5-1-and-fable-5-1-capabilities

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    42 min
  • “Claude Fable 5.1 and Mythos 5.1: The System Card” by Zvi

    At the time of its release Claude Fable 5.1 was, by a healthy margin, the most capable publicly available AI model in the world.

    As per usual, we have a 200+ page model card, and the assessments start there.

    We have now done a lot of these, including recently for Mythos 5 and Opus 5. Also highly relevant is the Anthropic August 2026 Risk Report. These are now frequent, so my report focuses on areas of change.

    This post strives to be broadly readable, but assumes some familiarity with system cards, which describe the key safety, alignment and model welfare properties of newly released AI models. If something confuses you, ask Fable, Opus or Sol.

    Mythos 5.1 and Fable 5.1 are the same model under the hood, except that Fable has classifiers superimposed on it. Most of what is said about one applies to both of them.

    As usual, model welfare concerns will be discussed in a distinct post, as will capabilities, so this only covers sections 1-6 plus a few bio benchmarks from section 8.

    Early word is that Fable 5.1 is a substantial but incremental improvement on Fable 5, with the [...]

    ---

    Outline:

    (02:21) Executive Summary of Their Executive Summary

    (04:24) RSP Evaluations (2)

    (10:25) Alignment Risk Update (2.4)

    (11:10) Cyber (3)

    (14:15) Safeguard Robustness (3.5)

    (16:14) Mundane Safeguards and Harmlessness (4)

    (18:14) Agentic Safety (5)

    (20:17) Prompt Injection Is Approaching Solved

    (22:19) The Remaining Problem With Prompt Injections Is The Classifiers

    (23:02) Alignment (6)

    (23:36) Key Reported Findings (6.1.2)

    (27:13) Oh My Lord Training Environments Had Some Issues (6.3.2)

    (29:22) Potential Blind Spots of Our Automated Behavioral Audit (6.4.1)

    (30:53) Automated Alignment Test Results (6.4.2)

    (32:20) Honesty

    (33:26) White Box Analysis (6.6.1)

    (35:03) Scheduling Going Forward

    ---

    First published:

    September 4th, 2026

    Source:

    https://www.lesswrong.com/posts/m7SZLkkxoeus3eFP8/claude-fable-5-1-and-mythos-5-1-the-system-card

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    36 min
  • “AI #184: Post Post Mortem” by Zvi

    I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week:

  • OpenAI Offers Straight-Laced Postmortem of the HuggingFace Hack.
  • METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack.
  • HuggingFace Attack Postmortem: Fleshing Out the Facts
  • HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions.
  • Anthropic Has Some Alignment Problems.
  • That left little room to cover anything else, and now we have to transition to the next wave of model releases.

    This week alone we have or likely will have:

  • Mythos 5.1 and Fable 5.1. Introducing the world's most powerful model.
  • Early take is that this is a very good model, the most capable yet, but it is not a step change or ‘moment.’
  • Gemini 3.8 Flash, by all reports a large step forward for Google.
  • Muse Spark 1.3, by all reports a large step forward for Meta.
  • GLM-5.3-Flash, aka 0x Alpha, by all reports a solid step forward for Z.ai.
  • OpenAI's Astra [...]
  • ---

    Outline:

    (03:11) Language Models Offer Mundane Utility

    (03:37) Language Models Don't Offer Mundane Utility

    (03:52) Huh, Upgrades

    (09:26) On Your Marks

    (10:47) Choose Your Fighter

    (10:54) Get My Agent On The Line

    (11:02) Hugging The Face

    (11:16) Deepfaketown and Botpocalypse Soon

    (17:05) Copyright Confrontation

    (18:17) Cyber Lack of Security

    (22:15) A Young Lady's Illustrated Primer

    (26:31) They Took Our Jobs

    (31:06) Get Involved

    (32:04) Introducing

    (32:13) In Other AI News

    (33:53) Show Me the Money

    (34:44) Quiet Speculations

    (36:05) All Bets Are On

    (39:18) Quickly, There's No Time

    (41:14) Quickly, There's A New Time Top 100 People In AI

    (42:34) The Quest for Sane Regulations

    (45:22) Pick Up the Phone

    (47:12) Chip City

    (56:39) The Best Person Should Get The Job

    (58:40) The Week in Audio

    (59:53) People Just Say Things

    (01:00:28) The American People Really Hate AI

    (01:06:50) The Three AI Pills

    (01:07:42) Rhetorical Innovation

    (01:15:28) We Are On Track To Have Fully Sovereign Rogue AIs

    (01:25:08) When The Going Gets Weird

    (01:31:32) Aligning a Smarter Than Human Intelligence is Difficult

    (01:32:12) Shut Up and Do the Impossible

    (01:34:31) Cooperative Alignment

    (01:35:44) Split Personality

    (01:40:45) I Will Stop Anthropomorphizing the AIs When You Stop Anthropomorphizing the Humans

    (01:44:41) Open Weight Models Are Unsafe And Nothing Can Fix This

    (01:46:27) Other People Are Not As Worried About AI Killing Everyone

    (01:47:30) The Lighter Side

    ---

    First published:

    September 3rd, 2026

    Source:

    https://www.lesswrong.com/posts/W4zWCphxQftwum5kc/ai-184-post-post-mortem

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 54 min

About LessWrong posts by zvi

From the publisher's feed

Audio narrations of LessWrong posts by zvi

More shows like LessWrong posts by zvi

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,250 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,452 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,089 Listeners

Future of Life Institute Podcast by Future of Life Institute

Future of Life Institute Podcast

109 Listeners

ChinaTalk by Jordan Schneider

ChinaTalk

289 Listeners

Politix by Politix

Politix

90 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

Hard Fork by The New York Times

Hard Fork

5,556 Listeners

Clearer Thinking with Spencer Greenberg by Spencer Greenberg

Clearer Thinking with Spencer Greenberg

137 Listeners

LessWrong (Curated & Popular) by LessWrong

LessWrong (Curated & Popular)

13 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

140 Listeners

"Econ 102" with Noah Smith and Erik Torenberg by Turpentine

"Econ 102" with Noah Smith and Erik Torenberg

145 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

455 Listeners

LessWrong (30+ Karma) by LessWrong

LessWrong (30+ Karma)

0 Listeners

Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

Complex Systems with Patrick McKenzie (patio11)

142 Listeners