
Sign up to save your podcasts
Or


What do I ultimately make of all the new versions of GPT-5?
The practical offerings and how they interact continues to change by the day. I expect more to come. It will take a while for things to settle down.
I’ll start with the central takeaways and how I select models right now, then go through the type and various questions in detail.
Table of Contents
Central Takeaways
My central takes [...]
---
Outline:
(00:33) Central Takeaways
(02:55) Choose Your Fighter
(05:43) Official Hype
(21:25) Chart Crime
(27:19) Model Crime
(28:09) Future Plans For OpenAI's Compute
(30:49) Rate Limitations
(32:19) The Routing Options Expand
(33:57) System Prompt
(35:14) On Writing
(41:01) Leading The Witness
(41:59) Hallucinations Are Down
(43:17) Best Of All Possible Worlds?
(47:53) Timelines
(58:25) Sycophancy Will Continue Because It Improves Morale
(01:00:22) Gaslighting Will Continue
(01:01:01) Going Pro
(01:04:09) Going Forward
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
A key problem with having and interpreting reactions to GPT-5 is that it is often unclear whether the reaction is to GPT-5, GPT-5-Router or GPT-5-Thinking.
Another is that many of the things people are reacting to changed rapidly after release, such as rate limits, the effectiveness of the model selection router and alternative options, and the availability of GPT-4o.
This complicates the tradition I have in new AI model reviews, which is to organize and present various representative and noteworthy reactions to the new model, to give a sense of what people are thinking and the diversity of opinion.
I also had make more cuts than usual, since there were so many eyes on this one. I tried to keep proportions similar to the original sample as best I could.
Reactions are organized roughly in order from positive to negative, with the drama around GPT-4o [...]
---
Outline:
(02:35) Tyler Cowen
(04:01) Ethan Mollick Thinks Ease Of Use Is A Big Deal
(05:42) The Router
(09:30) Remember To Use Thinking Mode
(14:50) The One Who Does Not Know How To Ask
(19:32) Nabeel Qureshi
(22:14) Other Positive Reactions
(27:01) It's A Good Model, Sir
(28:01) The Battle for Cursor Supremacy
(32:13) Automatic For The People
(36:17) Skeptical Reactions
(46:01) Colin Fraser Colin Frasiers
(48:13) I Want You Back
(01:01:09) The Verdict For Advanced Users Is Meh?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
GPT-5 was a long time coming.
Is it a good model, sir? Yes. In practice it is a good, but not great, model.
Or rather, it is several good models released at once: GPT-5, GPT-5-Thinking, GPT-5-With-The-Router, GPT-5-Pro, GPT-5-API. That leads to a lot of confusion.
What is most good? Cutting down on errors and hallucinations is a big deal. Ease of use and ‘just doing things’ have improved. Early reports are thinking mode is a large improvement on writing. Coding seems improved and can compete with Opus.
This first post covers an introduction, basic facts, benchmarks and the model card. Coverage will continue tomorrow.
This Fully Operational Battle Station
GPT-5 is here. They presented it as a really big deal. Death Star big.
Sam Altman (the night before release):
Nikita Bier: There is still time to delete.
PixelHulk:
Zvi [...]
---
Outline:
(01:04) This Fully Operational Battle Station
(04:20) Big Facts
(06:23) The System Card
(06:42) A Model By Any Other Name
(09:26) Safe Completions
(09:53) Mundane Safety
(10:48) Sycophancy
(14:46) The Art of the Jailbreak
(21:59) Hallucinations
(23:59) Deception
(27:58) Red Teaming
(29:03) Violent Attack Planning
(30:11) Prompt Injections
(32:20) Microsoft AI Red Teaming
(33:43) Preparedness Framework (Catastrophic and Existential Risks)
(33:49) Fine Tuning
(34:58) Safeguarding the API
(38:09) Biological Capabilities Remain Similar
(40:43) That One Graph From METR
(49:22) Big Compute
(49:53) On Your Marks
(57:00) Other People's Benchmarks
(01:01:22) Is That The Best You Can Do?
(01:03:08) Things To Come
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
---
Outline:
(01:15) Moderately Sized Models
(01:48) Introducing GPT-OSS
(03:56) The Model Card
(07:32) Our Price Cheap
(12:44) On Your Marks
(13:51) Mundane Safety Evaluations
(15:39) Preparedness Framework Evaluations
(21:03) Good Habits
(22:48) Distillation
(27:22) Safety First
(30:21) Other Reactions
(39:35) Hit Me Up I'm Open
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Brace for impact. We are presumably (checks watch) four hours from GPT-5.
That's the time you need to catch up on all the other AI news.
In another week, I might have done an entire post on Gemini 2.5 Deep Thinking, or Genie 3, or a few other things. This week? Quickly, there's no time.
OpenAI has already released an open model. I’m aiming to cover that tomorrow.
Table of Contents
Also: Claude 4.1 is an incremental improvement, On Altman's Interview With Theo Von.
---
Outline:
(00:42) Language Models Offer Mundane Utility
(03:04) Language Models Don't Offer Mundane Utility
(06:04) Huh, Upgrades
(07:31) On Your Marks
(11:24) Thinking Deeply With Gemini 2.5
(16:33) Choose Your Fighter
(16:47) Fun With Media Generation
(18:24) Optimal Optimization
(21:38) Get My Agent On The Line
(29:15) Deepfaketown and Botpocalypse Soon
(33:19) You Drive Me Crazy
(34:42) They Took Our Jobs
(37:53) Get Involved
(41:24) Introducing
(42:27) City In A Bottle
(48:15) Unprompted Suggestions
(48:54) In Other AI News
(51:10) Papers, Please
(51:18) The Mask Comes Off
(52:55) Show Me the Money
(01:01:11) Quiet Speculations
(01:09:20) Mark Zuckerberg Spreads Confusion
(01:15:05) The Quest for Sane Regulations
(01:19:11) David Sacks Once Again Amplifies Obvious Nonsense
(01:23:28) Chip City
(01:27:50) No Chip City
(01:30:00) Energy Crisis
(01:33:13) To The Moon
(01:34:58) Dario's Dismissal Deeply Disappoints, Depending on Details
(01:46:05) The Week in Audio
(01:50:48) Tyler Cowen Watch
(01:54:50) Rhetorical Innovation
(02:00:35) Shame Be Upon Them
(02:01:21) Correlation Causes Causation
(02:04:17) Aligning a Smarter Than Human Intelligence is Difficult
(02:06:29) The Lighter Side
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Claude Opus 4 has been updated to Claude Opus 4.1.
This is a correctly named incremental update, with the bigger news being ‘we plan to release substantially larger improvements to our models in the coming weeks.’
It is still worth noting if you code, as there are many indications this is a larger practical jump in performance than one might think.
We also got a change to the Claude.ai system prompt that helps with sycophancy and a few other issues, such as coming out and Saying The Thing more readily. It's going to be tricky to disentangle these changes, but that means Claude effectively got better for everyone, not only those doing agentic coding.
Tomorrow we get an OpenAI livestream that is presumably GPT-5, so I’m getting this out of the way now. Current plan is to cover GPT-OSS on Friday, and GPT-5 on Monday.
[...]---
Outline:
(01:01) Introducing Claude Opus 4.1
(05:25) The System Card
(09:56) Reactions
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
There's a time and a place for everything. It used to be called college.
Table of Contents
The Big Test
I am continuing to come around to the high-stakes-in-person-exam (or series of such exams) as the only practical solution to AI, also it was probably mostly the right answer already.
Sean T: It's [...]
---
Outline:
(00:16) The Big Test
(01:31) Testing, Testing
(03:44) Legalized Cheating On the Big Test
(06:12) What Happens When You Don't Test For Academics
(07:21) What Happens Without Academic Standards
(09:05) Another Academic Standard Perhaps
(11:07) RIP Columbia Core Curriculum and Also Social Theory
(14:15) College Tuition and Costs
(19:37) Negotiation
(20:45) Skipping College
(23:19) Respect Their Authoritah
(24:25) Men Skipping College
(28:05) Stanford Still Hates Fun
(30:35) Value of College
(31:10) Employment Prospects After College
(35:05) Fixing College
(36:17) Do Not Donate To A College
(38:43) Not Doing The Math
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Sam Altman talked recently to Theo Von.
Theo is genuinely engaging and curious throughout. This made me want to consider listening to his podcast more. I’d love to hang. He seems like a great dude.
The problem is that his curiosity has been redirected away from the places it would matter most – the Altman strategy of acting as if the biggest concerns, risks and problems flat out don’t exist successfully tricks Theo into not noticing them at all, and there are plenty of other things for him to focus on, so he does exactly that.
Meanwhile, Altman gets away with more of this ‘gentle singularity’ lie without using that term, letting it graduate to a background assumption. Dwarkesh would never.
Highlights, Quotes And Comments
Quotes are all from Altman.
Sam Altman: But also [kids born a [...]
---
Outline:
(00:56) Highlights, Quotes And Comments
(15:11) My Next Guest Needs No Introduction
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
There was enough governance related news this week to spin it out.
The EU AI Code of Practice
Anthropic, Google, OpenAI, Mistral, Aleph Alpha, Cohere and others commit to signing the EU AI Code of Practice. Google has now signed. Microsoft says it is likely to sign.
xAI signed the AI safety chapter of the code, but is refusing to sign the others, citing them as overreach especially as pertains to copyright.
The only company that said it would not sign at all is Meta.
This was the underreported story. All the important AI companies other than Meta have gotten behind the safety section of the EU AI Code of Practice. This represents a considerable strengthening of their commitments, and introduces an enforcement mechanism. Even Anthropic will be forced to step up parts of their game.
That leaves Meta as the rogue state [...]
---
Outline:
(00:13) The EU AI Code of Practice
(01:50) The Quest Against Regulations
(05:37) China Also Has An AI Action Plan
(15:07) Pick Up The Phone
(19:10) The AI Action Plan Has Good Marginal Proposals But Terrible Rhetoric
(29:44) Kratsios Explains The AI Action Plan
(39:22) Chip City
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Due to Continued Claude Code Complications, we can report Unlimited Usage Ultimately Unsustainable. May I suggest using the API, where Anthropic's yearly revenue is now projected to rise to $9 billion?
The biggest news items this week were in the policy realm, with the EU AI Code of Practice and the release of America's AI Action Plan and a Chinese response.
I am spinning off the policy realm into what is planned to be tomorrow's post (I’ve also spun off or pushed forward coverage of Altman's latest podcast, this time with Theo Von), so I’ll hit the highlights up here along with reviewing the week.
It turns out that when you focus on its concrete proposals, America's AI Action Plan Is Pretty Good. The people who wrote this knew what they were doing, and executed well given their world model and priorities. Most of the concrete [...]
---
Outline:
(02:17) Language Models Offer Mundane Utility
(06:36) Language Models Don't Offer Mundane Utility
(09:52) Huh, Upgrades
(10:57) Unlimited Usage Ultimately Unsustainable
(12:38) On Your Marks
(12:56) Are We Robot Or Are We Dancer
(13:41) Get My Agent On The Line
(14:54) Choose Your Fighter
(16:13) Code With Claude
(20:54) You Drive Me Crazy
(22:54) Deepfaketown and Botpocalypse Soon
(23:50) They Took Our Jobs
(29:52) Meta Promises Superglasses Or Something
(38:12) I Was Promised Flying Self-Driving Cars
(39:38) The Art of the Jailbreak
(40:05) Get Involved
(41:05) Introducing
(41:29) In Other AI News
(42:36) Show Me the Money
(47:53) Selling Out
(53:43) On Writing
(56:09) Quiet Speculations
(58:36) The Week in Audio
(59:42) Rhetorical Innovation
(01:07:57) Not Intentionally About AI
(01:08:55) Misaligned!
(01:10:27) Aligning A Dumber Than Human Intelligence Is Still Difficult
(01:11:06) Aligning a Smarter Than Human Intelligence is Difficult
(01:13:27) Subliminal Learning To Like The Owls
(01:19:38) Other People Are Not As Worried About AI Killing Everyone
(01:20:59) The Lighter Side
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
From the publisher's feed

26,250 Listeners

2,452 Listeners

1,089 Listeners

109 Listeners

289 Listeners

90 Listeners

572 Listeners

5,556 Listeners

137 Listeners

13 Listeners

140 Listeners

145 Listeners

455 Listeners

0 Listeners

142 Listeners