LessWrong posts by zvi

LessWrong posts by zvi

Download on the App Store

LessWrong posts by zvi episodes

  • “GPT-5s Are Alive: Synthesis” by Zvi

    What do I ultimately make of all the new versions of GPT-5?

    The practical offerings and how they interact continues to change by the day. I expect more to come. It will take a while for things to settle down.

    I’ll start with the central takeaways and how I select models right now, then go through the type and various questions in detail.

    Table of Contents

  • Central Takeaways.
  • Choose Your Fighter.
  • Official Hype.
  • Chart Crime.
  • Model Crime.
  • Future Plans For OpenAI's Compute.
  • Rate Limitations.
  • The Routing Options Expand.
  • System Prompt.
  • On Writing.
  • Leading The Witness.
  • Hallucinations Are Down.
  • Best Of All Possible Worlds?.
  • Timelines.
  • Sycophancy Will Continue Because It Improves Morale.
  • Gaslighting Will Continue.
  • Going Pro.
  • Going Forward.
  • Central Takeaways

    My central takes [...]

    ---

    Outline:

    (00:33) Central Takeaways

    (02:55) Choose Your Fighter

    (05:43) Official Hype

    (21:25) Chart Crime

    (27:19) Model Crime

    (28:09) Future Plans For OpenAI's Compute

    (30:49) Rate Limitations

    (32:19) The Routing Options Expand

    (33:57) System Prompt

    (35:14) On Writing

    (41:01) Leading The Witness

    (41:59) Hallucinations Are Down

    (43:17) Best Of All Possible Worlds?

    (47:53) Timelines

    (58:25) Sycophancy Will Continue Because It Improves Morale

    (01:00:22) Gaslighting Will Continue

    (01:01:01) Going Pro

    (01:04:09) Going Forward

    ---

    First published:

    August 13th, 2025

    Source:

    https://www.lesswrong.com/posts/4wYKkbkHooeQ4xznf/gpt-5s-are-alive-synthesis

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 6 min
  • “GPT-5s Are Alive: Outside Reactions, the Router and the Resurrection of GPT-4o” by Zvi

    A key problem with having and interpreting reactions to GPT-5 is that it is often unclear whether the reaction is to GPT-5, GPT-5-Router or GPT-5-Thinking.

    Another is that many of the things people are reacting to changed rapidly after release, such as rate limits, the effectiveness of the model selection router and alternative options, and the availability of GPT-4o.

    This complicates the tradition I have in new AI model reviews, which is to organize and present various representative and noteworthy reactions to the new model, to give a sense of what people are thinking and the diversity of opinion.

    I also had make more cuts than usual, since there were so many eyes on this one. I tried to keep proportions similar to the original sample as best I could.

    Reactions are organized roughly in order from positive to negative, with the drama around GPT-4o [...]

    ---

    Outline:

    (02:35) Tyler Cowen

    (04:01) Ethan Mollick Thinks Ease Of Use Is A Big Deal

    (05:42) The Router

    (09:30) Remember To Use Thinking Mode

    (14:50) The One Who Does Not Know How To Ask

    (19:32) Nabeel Qureshi

    (22:14) Other Positive Reactions

    (27:01) It's A Good Model, Sir

    (28:01) The Battle for Cursor Supremacy

    (32:13) Automatic For The People

    (36:17) Skeptical Reactions

    (46:01) Colin Fraser Colin Frasiers

    (48:13) I Want You Back

    (01:01:09) The Verdict For Advanced Users Is Meh?

    ---

    First published:

    August 12th, 2025

    Source:

    https://www.lesswrong.com/posts/uSGgByLKvRoKsDPih/gpt-5s-are-alive-outside-reactions-the-router-and-the

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 3 min
  • “GPT-5s Are Alive: Basic Facts, Benchmarks and the Model Card” by Zvi

    GPT-5 was a long time coming.

    Is it a good model, sir? Yes. In practice it is a good, but not great, model.

    Or rather, it is several good models released at once: GPT-5, GPT-5-Thinking, GPT-5-With-The-Router, GPT-5-Pro, GPT-5-API. That leads to a lot of confusion.

    What is most good? Cutting down on errors and hallucinations is a big deal. Ease of use and ‘just doing things’ have improved. Early reports are thinking mode is a large improvement on writing. Coding seems improved and can compete with Opus.

    This first post covers an introduction, basic facts, benchmarks and the model card. Coverage will continue tomorrow.

    This Fully Operational Battle Station

    GPT-5 is here. They presented it as a really big deal. Death Star big.

    Sam Altman (the night before release):

    Nikita Bier: There is still time to delete.

    PixelHulk:

    Zvi [...]

    ---

    Outline:

    (01:04) This Fully Operational Battle Station

    (04:20) Big Facts

    (06:23) The System Card

    (06:42) A Model By Any Other Name

    (09:26) Safe Completions

    (09:53) Mundane Safety

    (10:48) Sycophancy

    (14:46) The Art of the Jailbreak

    (21:59) Hallucinations

    (23:59) Deception

    (27:58) Red Teaming

    (29:03) Violent Attack Planning

    (30:11) Prompt Injections

    (32:20) Microsoft AI Red Teaming

    (33:43) Preparedness Framework (Catastrophic and Existential Risks)

    (33:49) Fine Tuning

    (34:58) Safeguarding the API

    (38:09) Biological Capabilities Remain Similar

    (40:43) That One Graph From METR

    (49:22) Big Compute

    (49:53) On Your Marks

    (57:00) Other People's Benchmarks

    (01:01:22) Is That The Best You Can Do?

    (01:03:08) Things To Come

    ---

    First published:

    August 11th, 2025

    Source:

    https://www.lesswrong.com/posts/4fLB2uzCcH6dEGnGs/gpt-5s-are-alive-basic-facts-benchmarks-and-the-model-card

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 5 min
  • “OpenAI’s GPT-OSS Is Already Old News” by Zvi
    That's on OpenAI. I don’t schedule their product releases.
    Since it takes several days to gather my reports on new models, we are doing our coverage of the OpenAI open weights models, GPT-OSS-20b and GPT-OSS-120b, today, after the release of GPT-5.
    The bottom line is that they seem like clearly good models in their targeted reasoning domains. There are many reports of them struggling in other domains, including with tool use, and they have very little inherent world knowledge, and the safety mechanisms appear obtrusive enough that many are complaining. It's not clear what they will be used for other than distillation into Chinese models.
    It is hard to tell, because open weight models need to be configured properly, and there are reports that many are doing this wrong, which could lead to clouded impressions. We will want to check back in a bit.
    In the Substack version of this [...]

    ---

    Outline:

    (01:15) Moderately Sized Models

    (01:48) Introducing GPT-OSS

    (03:56) The Model Card

    (07:32) Our Price Cheap

    (12:44) On Your Marks

    (13:51) Mundane Safety Evaluations

    (15:39) Preparedness Framework Evaluations

    (21:03) Good Habits

    (22:48) Distillation

    (27:22) Safety First

    (30:21) Other Reactions

    (39:35) Hit Me Up I'm Open

    ---

    First published:

    August 8th, 2025

    Source:

    https://www.lesswrong.com/posts/AJ94X73M6KgAZFJH2/openai-s-gpt-oss-is-already-old-news

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    42 min
  • “AI #128: Four Hours Until Probably Not The Apocalypse” by Zvi

    Brace for impact. We are presumably (checks watch) four hours from GPT-5.

    That's the time you need to catch up on all the other AI news.

    In another week, I might have done an entire post on Gemini 2.5 Deep Thinking, or Genie 3, or a few other things. This week? Quickly, there's no time.

    OpenAI has already released an open model. I’m aiming to cover that tomorrow.

    Table of Contents

    Also: Claude 4.1 is an incremental improvement, On Altman's Interview With Theo Von.

  • Language Models Offer Mundane Utility. The only help you need?
  • Language Models Don’t Offer Mundane Utility. Can’t use Claude to train GPT-5.
  • Huh, Upgrades. ChatGPT for government, Gemini for students, Claude security.
  • On Your Marks. More analysis of Psyho's victory over OpenAI at AWTF.
  • Thinking Deeply With Gemini 2.5. The power of parallel [...]
  • ---

    Outline:

    (00:42) Language Models Offer Mundane Utility

    (03:04) Language Models Don't Offer Mundane Utility

    (06:04) Huh, Upgrades

    (07:31) On Your Marks

    (11:24) Thinking Deeply With Gemini 2.5

    (16:33) Choose Your Fighter

    (16:47) Fun With Media Generation

    (18:24) Optimal Optimization

    (21:38) Get My Agent On The Line

    (29:15) Deepfaketown and Botpocalypse Soon

    (33:19) You Drive Me Crazy

    (34:42) They Took Our Jobs

    (37:53) Get Involved

    (41:24) Introducing

    (42:27) City In A Bottle

    (48:15) Unprompted Suggestions

    (48:54) In Other AI News

    (51:10) Papers, Please

    (51:18) The Mask Comes Off

    (52:55) Show Me the Money

    (01:01:11) Quiet Speculations

    (01:09:20) Mark Zuckerberg Spreads Confusion

    (01:15:05) The Quest for Sane Regulations

    (01:19:11) David Sacks Once Again Amplifies Obvious Nonsense

    (01:23:28) Chip City

    (01:27:50) No Chip City

    (01:30:00) Energy Crisis

    (01:33:13) To The Moon

    (01:34:58) Dario's Dismissal Deeply Disappoints, Depending on Details

    (01:46:05) The Week in Audio

    (01:50:48) Tyler Cowen Watch

    (01:54:50) Rhetorical Innovation

    (02:00:35) Shame Be Upon Them

    (02:01:21) Correlation Causes Causation

    (02:04:17) Aligning a Smarter Than Human Intelligence is Difficult

    (02:06:29) The Lighter Side

    ---

    First published:

    August 7th, 2025

    Source:

    https://www.lesswrong.com/posts/DSxnqyok2g4NaMp8u/ai-128-four-hours-until-probably-not-the-apocalypse

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    2 hr 9 min
  • “Opus 4.1 Is An Incremental Improvement” by Zvi

    Claude Opus 4 has been updated to Claude Opus 4.1.

    This is a correctly named incremental update, with the bigger news being ‘we plan to release substantially larger improvements to our models in the coming weeks.’

    It is still worth noting if you code, as there are many indications this is a larger practical jump in performance than one might think.

    We also got a change to the Claude.ai system prompt that helps with sycophancy and a few other issues, such as coming out and Saying The Thing more readily. It's going to be tricky to disentangle these changes, but that means Claude effectively got better for everyone, not only those doing agentic coding.

    Tomorrow we get an OpenAI livestream that is presumably GPT-5, so I’m getting this out of the way now. Current plan is to cover GPT-OSS on Friday, and GPT-5 on Monday.

    [...]

    ---

    Outline:

    (01:01) Introducing Claude Opus 4.1

    (05:25) The System Card

    (09:56) Reactions

    ---

    First published:

    August 6th, 2025

    Source:

    https://www.lesswrong.com/posts/hicuZJQwRYCiFCZbq/opus-4-1-is-an-incremental-improvement

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    14 min
  • “Childhood and Education #13: College” by Zvi

    There's a time and a place for everything. It used to be called college.

    Table of Contents

  • The Big Test.
  • Testing, Testing.
  • Legalized Cheating On the Big Test.
  • What Happens When You Don’t Test For Academics.
  • What Happens Without Academic Standards.
  • Another Academic Standard Perhaps.
  • RIP Columbia Core Curriculum and Also Social Theory.
  • College Tuition and Costs.
  • Negotiation.
  • Skipping College.
  • Respect Their Authoritah.
  • Men Skipping College.
  • Stanford Still Hates Fun.
  • Value of College.
  • Employment Prospects After College.
  • Fixing College.
  • Do Not Donate To A College.
  • Not Doing The Math.
  • The Big Test

    I am continuing to come around to the high-stakes-in-person-exam (or series of such exams) as the only practical solution to AI, also it was probably mostly the right answer already.

    Sean T: It's [...]

    ---

    Outline:

    (00:16) The Big Test

    (01:31) Testing, Testing

    (03:44) Legalized Cheating On the Big Test

    (06:12) What Happens When You Don't Test For Academics

    (07:21) What Happens Without Academic Standards

    (09:05) Another Academic Standard Perhaps

    (11:07) RIP Columbia Core Curriculum and Also Social Theory

    (14:15) College Tuition and Costs

    (19:37) Negotiation

    (20:45) Skipping College

    (23:19) Respect Their Authoritah

    (24:25) Men Skipping College

    (28:05) Stanford Still Hates Fun

    (30:35) Value of College

    (31:10) Employment Prospects After College

    (35:05) Fixing College

    (36:17) Do Not Donate To A College

    (38:43) Not Doing The Math

    ---

    First published:

    August 5th, 2025

    Source:

    https://www.lesswrong.com/posts/WzN3PFpyBaMeXqeRz/childhood-and-education-13-college

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    45 min
  • “On Altman’s Interview With Theo Von” by Zvi

    Sam Altman talked recently to Theo Von.

    Double click to interact with video

    Theo is genuinely engaging and curious throughout. This made me want to consider listening to his podcast more. I’d love to hang. He seems like a great dude.

    The problem is that his curiosity has been redirected away from the places it would matter most – the Altman strategy of acting as if the biggest concerns, risks and problems flat out don’t exist successfully tricks Theo into not noticing them at all, and there are plenty of other things for him to focus on, so he does exactly that.

    Meanwhile, Altman gets away with more of this ‘gentle singularity’ lie without using that term, letting it graduate to a background assumption. Dwarkesh would never.

    Highlights, Quotes And Comments

    Quotes are all from Altman.

    Sam Altman: But also [kids born a [...]

    ---

    Outline:

    (00:56) Highlights, Quotes And Comments

    (15:11) My Next Guest Needs No Introduction

    ---

    First published:

    August 4th, 2025

    Source:

    https://www.lesswrong.com/posts/XXfi7rHgjki8RxxyL/on-altman-s-interview-with-theo-von

    ---

    Narrated by TYPE III AUDIO.

    17 min
  • “The Week in AI Governance” by Zvi

    There was enough governance related news this week to spin it out.

    The EU AI Code of Practice

    Anthropic, Google, OpenAI, Mistral, Aleph Alpha, Cohere and others commit to signing the EU AI Code of Practice. Google has now signed. Microsoft says it is likely to sign.

    xAI signed the AI safety chapter of the code, but is refusing to sign the others, citing them as overreach especially as pertains to copyright.

    The only company that said it would not sign at all is Meta.

    This was the underreported story. All the important AI companies other than Meta have gotten behind the safety section of the EU AI Code of Practice. This represents a considerable strengthening of their commitments, and introduces an enforcement mechanism. Even Anthropic will be forced to step up parts of their game.

    That leaves Meta as the rogue state [...]

    ---

    Outline:

    (00:13) The EU AI Code of Practice

    (01:50) The Quest Against Regulations

    (05:37) China Also Has An AI Action Plan

    (15:07) Pick Up The Phone

    (19:10) The AI Action Plan Has Good Marginal Proposals But Terrible Rhetoric

    (29:44) Kratsios Explains The AI Action Plan

    (39:22) Chip City

    ---

    First published:

    August 1st, 2025

    Source:

    https://www.lesswrong.com/posts/9rqMPLdpctxig2iAg/the-week-in-ai-governance

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    47 min
  • “AI #127: Continued Claude Code Complications” by Zvi

    Due to Continued Claude Code Complications, we can report Unlimited Usage Ultimately Unsustainable. May I suggest using the API, where Anthropic's yearly revenue is now projected to rise to $9 billion?

    The biggest news items this week were in the policy realm, with the EU AI Code of Practice and the release of America's AI Action Plan and a Chinese response.

    I am spinning off the policy realm into what is planned to be tomorrow's post (I’ve also spun off or pushed forward coverage of Altman's latest podcast, this time with Theo Von), so I’ll hit the highlights up here along with reviewing the week.

    It turns out that when you focus on its concrete proposals, America's AI Action Plan Is Pretty Good. The people who wrote this knew what they were doing, and executed well given their world model and priorities. Most of the concrete [...]

    ---

    Outline:

    (02:17) Language Models Offer Mundane Utility

    (06:36) Language Models Don't Offer Mundane Utility

    (09:52) Huh, Upgrades

    (10:57) Unlimited Usage Ultimately Unsustainable

    (12:38) On Your Marks

    (12:56) Are We Robot Or Are We Dancer

    (13:41) Get My Agent On The Line

    (14:54) Choose Your Fighter

    (16:13) Code With Claude

    (20:54) You Drive Me Crazy

    (22:54) Deepfaketown and Botpocalypse Soon

    (23:50) They Took Our Jobs

    (29:52) Meta Promises Superglasses Or Something

    (38:12) I Was Promised Flying Self-Driving Cars

    (39:38) The Art of the Jailbreak

    (40:05) Get Involved

    (41:05) Introducing

    (41:29) In Other AI News

    (42:36) Show Me the Money

    (47:53) Selling Out

    (53:43) On Writing

    (56:09) Quiet Speculations

    (58:36) The Week in Audio

    (59:42) Rhetorical Innovation

    (01:07:57) Not Intentionally About AI

    (01:08:55) Misaligned!

    (01:10:27) Aligning A Dumber Than Human Intelligence Is Still Difficult

    (01:11:06) Aligning a Smarter Than Human Intelligence is Difficult

    (01:13:27) Subliminal Learning To Like The Owls

    (01:19:38) Other People Are Not As Worried About AI Killing Everyone

    (01:20:59) The Lighter Side

    ---

    First published:

    July 31st, 2025

    Source:

    https://www.lesswrong.com/posts/GeTssgFDCvHDGAvhT/ai-127-continued-claude-code-complications

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 24 min

About LessWrong posts by zvi

From the publisher's feed

Audio narrations of LessWrong posts by zvi

More shows like LessWrong posts by zvi

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,250 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,452 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,089 Listeners

Future of Life Institute Podcast by Future of Life Institute

Future of Life Institute Podcast

109 Listeners

ChinaTalk by Jordan Schneider

ChinaTalk

289 Listeners

Politix by Politix

Politix

90 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

Hard Fork by The New York Times

Hard Fork

5,556 Listeners

Clearer Thinking with Spencer Greenberg by Spencer Greenberg

Clearer Thinking with Spencer Greenberg

137 Listeners

LessWrong (Curated & Popular) by LessWrong

LessWrong (Curated & Popular)

13 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

140 Listeners

"Econ 102" with Noah Smith and Erik Torenberg by Turpentine

"Econ 102" with Noah Smith and Erik Torenberg

145 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

455 Listeners

LessWrong (30+ Karma) by LessWrong

LessWrong (30+ Karma)

0 Listeners

Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

Complex Systems with Patrick McKenzie (patio11)

142 Listeners