LessWrong posts by zvi

LessWrong posts by zvi

Download on the App Store

LessWrong posts by zvi episodes

  • “Kimi K2” by Zvi

    While most people focused on Grok, there was another model release that got uniformly high praise: Kimi K2 from Moonshot.ai.

    It's definitely a good model, sir, especially for a cheap-to-run open model.

    It is plausibly the best model for creative writing, outright. It is refreshingly different, and opens up various doors through which one can play. And it proves the value of its new architecture.

    It is not an overall SoTA frontier model, but it is not trying to be one.

    The reasoning model version is coming. Price that in now.

    Introducing Kimi K2

    Introducing the latest model that matters, Kimi K2.

    Hello, Kimi K2! Open-Source Agentic Model!

    1T total / 32B active MoE model

    SOTA on SWE Bench Verified, Tau2 & AceBench among open models

    Strong in coding and agentic tasks

    Multimodal & thought-mode not supported for [...]

    ---

    Outline:

    (00:45) Introducing Kimi K2

    (02:24) Having a Moment

    (03:29) Another Nimble Effort

    (05:37) On Your Marks

    (07:48) Everybody Loves Kimi, Baby

    (13:09) Okay, Not Quite Everyone

    (14:06) Everyone Uses Kimi, Baby

    (15:42) Write Like A Human

    (25:32) What Happens Next

    ---

    First published:

    July 16th, 2025

    Source:

    https://www.lesswrong.com/posts/qsyj37hwh9N8kcopJ/kimi-k2

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    27 min
  • “Grok 4 Various Things” by Zvi

    Yesterday I covered a few rather important Grok incidents.

    Today is all about Grok 4's capabilities and features. Is it a good model, sir?

    It's not a great model. It's not the smartest or best model.

    But it's at least an okay model. Probably a ‘good’ model.

    Talking a Big Game

    xAI was given a goal. They were to release something that could, ideally with a straight face, be called ‘the world's smartest artificial intelligence.’

    On that level, well, congratulations to Elon Musk and xAI. You have successfully found benchmarks that enable you to make that claim.

    xAI: We just unveiled Grok 4, the world's smartest artificial intelligence.

    Grok 4 outperforms all other models on the ARC-AGI benchmark, scoring 15.9% – nearly double that of the next best model – and establishing itself as the most intelligent AI to date.

    [...]

    ---

    Outline:

    (00:30) Talking a Big Game

    (03:57) Gotta Go Fast

    (04:38) On Your Marks

    (07:21) Some Key Facts About Grok 4

    (09:44) SuperGrok Heavy, Man

    (11:43) Blunt Instrument

    (13:40) Easiest Jailbreak Ever

    (15:49) ARC-AGI-2

    (17:08) Gaming the Benchmarks

    (21:23) Why Care About Benchmarks?

    (23:29) Other People's Benchmarks

    (32:14) Impressed Reactions to Grok

    (37:47) Coding Specific Feedback

    (40:00) Unimpressed Reactions to Grok

    (46:39) Tyler Cowen Is Not Impressed

    (48:26) You Had One Job

    (49:38) Reactions to Reactions Overall

    (51:36) The MechaHitler Lives On

    (53:03) But Wait, There's More

    (01:00:12) Sixth Law Of Human Stupidity Strikes Again

    (01:05:09) There I Fixed It

    (01:06:57) What Is Grok 4 And What Should We Make Of It?

    ---

    First published:

    July 15th, 2025

    Source:

    https://www.lesswrong.com/posts/ciuKn9aktXxJ2K6Rc/grok-4-various-things

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 14 min
  • “Worse Than MechaHitler” by Zvi

    Grok 4, which has excellent benchmarks and which xAI claims is ‘the world's smartest artificial intelligence,’ is the big news.

    If you set aside the constant need to say ‘No, Grok, No,’ is it a good model, sir?

    My take in terms of its capabilities, which I will expand upon at great length later this week: It is a good model. Not a great model. Not the best model. Not ‘the world's smartest artificial intelligence.’ There do not seem to be any great use cases to choose it over alternatives, unless you are searching Twitter. But it is a good model.

    There is a catch. There are many reasons one might not want to trust it, on a different level than the reasons not to trust models from other labs. There has been a series of epic failures and poor choices, which will be difficult to [...]

    ---

    Outline:

    (01:33) The System Prompt

    (08:07) MechaHitler

    (10:17) The Official Explanation of MechaHitler

    (17:53) Worse Than MechaHitler

    (22:22) Unintended Behavior

    (26:22) Off Based

    (28:37) Your Face Will Be Stuck That Way

    (30:24) I Couldn't Do Solve Problem In Several Hours So It Must Be Very Hard

    (38:57) Safety Third

    (44:25) How Bad Are Things?

    ---

    First published:

    July 14th, 2025

    Source:

    https://www.lesswrong.com/posts/YmdCN5GBwkud5ZzYx/worse-than-mechahitler

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    46 min
  • “OpenAI Model Differentiation 101” by Zvi

    LLMs can be deeply confusing. Thanks to a commission, today we go back to basics.

    How did we get such a wide array of confusingly named and labeled models and modes in ChatGPT? What are they, and when and why would you use each of them for what purposes, and how does this relate to what is available elsewhere? How does this relate to hallucinations, sycophancy and other basic issues, and what are the basic ways of mitigating those issues?

    If you already know these basics, you can and should skip this post.

    This is a reference, and a guide for the new and the perplexed, until the time comes that they change everything again, presumably with GPT-5.

    A Brief History of OpenAI Models and Their Names

    Tech companies are notorious for being terrible at naming things. One decision that seems like the best [...]

    ---

    Outline:

    (00:51) A Brief History of OpenAI Models and Their Names

    (06:05) The Models We Have Now in ChatGPT

    (12:23) What About The Competition?

    (12:51) Claude (Claude.ai)

    (14:30) Gemini

    (16:03) Grok

    (16:59) Hallucinations

    (19:09) Sycophancy

    (20:12) Going Beyond

    ---

    First published:

    July 11th, 2025

    Source:

    https://www.lesswrong.com/posts/5NF7DRvcLLGHn78bT/openai-model-differentiation-101

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    22 min
  • “AI #124: Grokless Interlude” by Zvi
    Last night, on the heels of some rather unfortunate incidents involving the Twitter version of Grok 3, xAI released Grok 4. There are some impressive claimed benchmarks. As per usual, I will wait a few days so others can check it out, and then offer my take early next week, and this post otherwise won’t discuss Grok 4 further.
    There are plenty of other things to look into while we wait for that.
    I am also not yet covering Anthropic's latest alignment faking paper, which may well get its own post.

    Table of Contents

  • Language Models Offer Mundane Utility. Who is 10x more productive?
  • Language Models Don’t Offer Mundane Utility. Branching paths.
  • Huh, Upgrades. DR in the OAI API, plus a tool called Study Together.
  • Preserve Our History. What are the barriers to availability of Opus 3?
  • Choose Your Fighter. GPT-4o offers [...]
  • ---

    Outline:

    (00:43) Language Models Offer Mundane Utility

    (05:08) Language Models Don't Offer Mundane Utility

    (06:57) Huh, Upgrades

    (07:53) Preserve Our History

    (11:18) Choose Your Fighter

    (12:36) Wouldn't You Prefer A Good Game of Chess

    (14:30) Fun With Media Generation

    (14:40) No Grok No

    (16:29) Deepfaketown and Botpocalypse Soon

    (19:15) Unprompted Attention

    (20:11) Overcoming Bias

    (22:18) Get My Agent On The Line

    (23:40) They Took Our Jobs

    (27:59) Get Involved

    (28:27) Introducing

    (30:11) In Other AI News

    (32:28) Show Me the Money

    (34:59) The Explanation Is Always Transaction Costs

    (37:56) Quiet Speculations

    (44:23) Genesis

    (46:29) The Quest for Sane Regulations

    (52:06) Chip City

    (52:28) Choosing The Right Regulatory Target

    (01:00:42) The Week in Audio

    (01:01:15) Rhetorical Innovation

    (01:04:33) Aligning a Smarter Than Human Intelligence is Difficult

    (01:10:10) Don't Worry We Have Human Oversight

    (01:14:09) Don't Worry We Have Chain Of Thought Monitoring

    (01:18:47) Sycophancy Is Hard To Fix

    (01:21:43) The Lighter Side

    ---

    First published:

    July 10th, 2025

    Source:

    https://www.lesswrong.com/posts/FczrW2kQ7WxGW39Yv/ai-124-grokless-interlude

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 24 min
  • “No, Grok, No” by Zvi

    It was the July 4 weekend. Grok on Twitter got some sort of upgrade.

    Elon Musk: We have improved @Grok significantly.

    You should notice a difference when you ask Grok questions.

    Indeed we did notice big differences.

    It did not go great. Then it got worse.

    That does not mean low quality answers or being a bit politically biased. Nor does it mean one particular absurd quirk like we saw in Regarding South Africa, or before that the narrow instruction not to criticize particular individuals.

    Here ‘got worse’ means things that involve the term ‘MechaHitler.’

    Doug Borton: I did Nazi this coming.

    Perhaps we should have. Three (escalating) times is enemy action.

    I had very low expectations for xAI, including on these topics. But not like this.

    In the wake of these events, Linda Yaccarino has stepped down this [...]

    ---

    Outline:

    (01:29) Finger On The Scale

    (05:06) We Got Trouble

    (07:52) Finger Somewhere Else

    (09:32) Worst Of The Worst

    (11:16) Fun Messing With Grok

    (14:06) The Hitler Coefficient

    (20:20) MechaHitler

    (21:42) The Two Groks

    (22:41) I'm Shocked, Shocked, Well Not Shocked

    (24:05) Misaligned!

    (31:39) Nothing To See Here

    (33:17) He Just Tweeted It Out

    (36:05) What Have We Learned?

    ---

    First published:

    July 9th, 2025

    Source:

    https://www.lesswrong.com/posts/CE8W4GEofRwHe4fiu/no-grok-no

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    39 min
  • “Balsa Update: Springtime in DC” by Zvi
    Today's post is an update from my contractor at Balsa Research, Jennifer Chen. I offer guidance and make strategic choices, but she's the one who makes the place run. Among all the other crazy things that have been happening lately, we had to divert some time from our Jones Act efforts to fight against some potentially far more disastrous regulations that got remarkably close to happening.

    Springtime in DC for Balsa, by Jennifer Chen

    What if, in addition to restricting all domestic waterborne trade to U.S.-built, U.S-flagged vessels, we also required the same of 20% of all U.S. exports?
    In late February this year, Balsa Research got word that this was a serious new proposal coming out of the USTR, with public comments due soon and public hearings not much longer after that.
    The glaring problem with this proposal was that there were fewer than one hundred oceangoing ships [...]

    ---

    Outline:

    (00:33) Springtime in DC for Balsa, by Jennifer Chen

    (03:35) Why was no one else talking about this?

    (07:27) How Likely Did We Think This Was Going to be Enacted?

    (08:50) Okay, Balsa Should Do Something

    (10:07) Balsa at the USTR Public Hearings

    (12:34) Did Balsa... Do Anything?

    (13:50) Should Balsa Continue to Do Things?

    (15:26) Balsa Research is Once More 100% Focused on Jones Act Reform

    ---

    First published:

    July 8th, 2025

    Source:

    https://www.lesswrong.com/posts/tiwjxgGSxSysHzAuc/balsa-update-springtime-in-dc

    ---

    Narrated by TYPE III AUDIO.

    20 min
  • “On Alpha School” by Zvi

    The epic 18k word writeup on Austin's flagship Alpha School is excellent. It is long, but given the blog you’re reading now, if you have interest in such topics I’d strongly consider reading the whole thing.

    One must always take such claims and reports with copious salt. But in terms of the core claims about what is happening and why it is happening, I find this mostly credible. I don’t know how far it can scale but I suspect quite far. None of this involves anything surprising, and none of it even involves much use of generative AI.

    Rui Ma here gives a shorter summary and offers takeaways compatible with mine.

    Table of Contents

  • What Is It?
  • What It Isn’t.
  • Intrinsic Versus Extrinsic Motivation.
  • High Versus Low Structure Learners.
  • I’ve Got a Theory.
  • Is This Really The True Objection?
  • [...]

    ---

    Outline:

    (00:47) What Is It?

    (05:00) What It Isn't

    (08:56) Intrinsic Versus Extrinsic Motivation

    (13:30) High Versus Low Structure Learners

    (14:06) I've Got a Theory

    (17:54) Is This Really The True Objection?

    ---

    First published:

    July 7th, 2025

    Source:

    https://www.lesswrong.com/posts/vwNygY4puHunjv6Pk/on-alpha-school

    ---

    Narrated by TYPE III AUDIO.

    25 min
  • “Housing Roundup #12” by Zvi

    Abundance and YIMBY are on the march. Things are looking good. The wins are each small, but every little bit helps. There are lots of different little things you can do. In theory you have to worry about a homeostatic model where solving some problems causes locals to double down on other barriers, but this seems to not be what we see.

    There are definitely important exceptions. Los Angeles is not so interested in rebuilding from the fires and backpaddled the moment developers started to actually build 100% affordable housing because somehow that was a bad thing. New York's democratic party nominated who they nominated. Massachusetts wants to seal eviction records.

    Overall, though, it's hard not to be hopeful right now. Even when we see bad policies, they are couched increasingly in the rhetoric of good goals and policies. In the long term, that leads to wins.

    [...]

    ---

    Outline:

    (01:11) Rent Control

    (02:52) Affordable Housing

    (09:19) A Vision

    (09:56) Private Equity

    (11:14) Home for Rent

    (14:08) Making Housing Worse On Purpose So You Can Click

    (16:47) Open Philanthropy Strikes Again

    (18:08) The Abundance Debate

    (21:24) Single Staircase Apartment Buildings

    (25:18) Dublin

    (25:43) Western Housing Costs

    (27:22) Los Angeles

    (28:16) LA Fire

    (31:15) San Francisco

    (34:36) California

    (39:27) Oregon

    (41:34) Montana

    (43:58) Maine

    (44:50) North Carolina

    (45:10) New York City

    (49:10) Massachusetts

    (51:43) Texas

    (54:04) Poland

    ---

    First published:

    July 4th, 2025

    Source:

    https://www.lesswrong.com/posts/wuoTsXoe93mXavofB/housing-roundup-12

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    56 min
  • “AI #123: Moratorium Moratorium” by Zvi

    The big AI story this week was the battle over the insane AI regulatory moratorium, which came dangerously close to passing. Ultimately, after Senator Blackburn realized her deal was no good and backed out of it, the dam broke, and ultimately the Senate voted 99-1 to strip the moratorium out of the BBB. I also covered last week's hopeful house hearing in detail, so we can remember this as a reference point.

    Otherwise, plenty of other things happened but in terms of big items this was a relatively quiet week. Always enjoy such respites while they last. Next week we are told we are getting Grok 4.

    Table of Contents

  • Table of Contents.
  • Language Models Offer Mundane Utility. Submit to the Amanda Askell hypnosis.
  • It Is I, Claudius, Vender of Items. Claude tries to run a vending machine.
  • Language Models Don’t Offer [...]
  • ---

    Outline:

    (00:50) Language Models Offer Mundane Utility

    (02:39) It Is I, Claudius, Vender of Items

    (05:36) Language Models Don't Offer Mundane Utility

    (06:10) GPT-4o Is An Absurd Sycophant

    (06:50) Preserve Our History

    (08:03) Fork In The Road

    (09:43) On Your Marks

    (11:06) Choose Your Fighter

    (12:32) Deepfaketown and Botpocalypse Soon

    (15:19) Goodhart's Law Strikes Again

    (18:51) Get My Agent On The Line

    (21:09) They Took Our Jobs

    (22:36) Get Involved

    (23:50) Introducing

    (25:25) Copyright Confrontation

    (27:00) Show Me the Money

    (33:15) Quiet Speculations

    (34:29) Minimum Viable Model

    (39:14) Timelines

    (41:25) Considering Chilling Out

    (46:17) The Quest for Sane Regulations

    (51:28) The Committee Recommends

    (58:00) Chip City

    (01:00:53) The Week in Audio

    (01:05:33) Rhetorical Innovation

    (01:15:53) Please Speak Directly Into The Microphone

    (01:17:08) Gary Marcus Predicts

    (01:22:10) The Vibes They Are A-Changing

    (01:27:03) Misaligned!

    (01:28:52) Aligning a Smarter Than Human Intelligence is Difficult

    (01:32:45) The Lighter Side

    The original text contained 1 footnote which was omitted from this narration.

    ---

    First published:

    July 3rd, 2025

    Source:

    https://www.lesswrong.com/posts/9bbu9nebdSfXa8cKa/ai-123-moratorium-moratorium

    ---

    Narrated by TYPE III AUDIO.

    ---

    Images from the article:

    Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

    1 hr 35 min

About LessWrong posts by zvi

From the publisher's feed

Audio narrations of LessWrong posts by zvi

More shows like LessWrong posts by zvi

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,250 Listeners

Conversations with Tyler by Mercatus Center at George Mason University

Conversations with Tyler

2,452 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,089 Listeners

Future of Life Institute Podcast by Future of Life Institute

Future of Life Institute Podcast

109 Listeners

ChinaTalk by Jordan Schneider

ChinaTalk

289 Listeners

Politix by Politix

Politix

90 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

572 Listeners

Hard Fork by The New York Times

Hard Fork

5,556 Listeners

Clearer Thinking with Spencer Greenberg by Spencer Greenberg

Clearer Thinking with Spencer Greenberg

137 Listeners

LessWrong (Curated & Popular) by LessWrong

LessWrong (Curated & Popular)

13 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

140 Listeners

"Econ 102" with Noah Smith and Erik Torenberg by Turpentine

"Econ 102" with Noah Smith and Erik Torenberg

145 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

455 Listeners

LessWrong (30+ Karma) by LessWrong

LessWrong (30+ Karma)

0 Listeners

Complex Systems with Patrick McKenzie (patio11) by Patrick McKenzie

Complex Systems with Patrick McKenzie (patio11)

142 Listeners