
Sign up to save your podcasts
Or


While most people focused on Grok, there was another model release that got uniformly high praise: Kimi K2 from Moonshot.ai.
It's definitely a good model, sir, especially for a cheap-to-run open model.
It is plausibly the best model for creative writing, outright. It is refreshingly different, and opens up various doors through which one can play. And it proves the value of its new architecture.
It is not an overall SoTA frontier model, but it is not trying to be one.
The reasoning model version is coming. Price that in now.
Introducing Kimi K2
Introducing the latest model that matters, Kimi K2.
Hello, Kimi K2! Open-Source Agentic Model!
1T total / 32B active MoE model
SOTA on SWE Bench Verified, Tau2 & AceBench among open models
Strong in coding and agentic tasks
Multimodal & thought-mode not supported for [...]
---
Outline:
(00:45) Introducing Kimi K2
(02:24) Having a Moment
(03:29) Another Nimble Effort
(05:37) On Your Marks
(07:48) Everybody Loves Kimi, Baby
(13:09) Okay, Not Quite Everyone
(14:06) Everyone Uses Kimi, Baby
(15:42) Write Like A Human
(25:32) What Happens Next
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Yesterday I covered a few rather important Grok incidents.
Today is all about Grok 4's capabilities and features. Is it a good model, sir?
It's not a great model. It's not the smartest or best model.
But it's at least an okay model. Probably a ‘good’ model.
Talking a Big Game
xAI was given a goal. They were to release something that could, ideally with a straight face, be called ‘the world's smartest artificial intelligence.’
On that level, well, congratulations to Elon Musk and xAI. You have successfully found benchmarks that enable you to make that claim.
xAI: We just unveiled Grok 4, the world's smartest artificial intelligence.
Grok 4 outperforms all other models on the ARC-AGI benchmark, scoring 15.9% – nearly double that of the next best model – and establishing itself as the most intelligent AI to date.
[...]
---
Outline:
(00:30) Talking a Big Game
(03:57) Gotta Go Fast
(04:38) On Your Marks
(07:21) Some Key Facts About Grok 4
(09:44) SuperGrok Heavy, Man
(11:43) Blunt Instrument
(13:40) Easiest Jailbreak Ever
(15:49) ARC-AGI-2
(17:08) Gaming the Benchmarks
(21:23) Why Care About Benchmarks?
(23:29) Other People's Benchmarks
(32:14) Impressed Reactions to Grok
(37:47) Coding Specific Feedback
(40:00) Unimpressed Reactions to Grok
(46:39) Tyler Cowen Is Not Impressed
(48:26) You Had One Job
(49:38) Reactions to Reactions Overall
(51:36) The MechaHitler Lives On
(53:03) But Wait, There's More
(01:00:12) Sixth Law Of Human Stupidity Strikes Again
(01:05:09) There I Fixed It
(01:06:57) What Is Grok 4 And What Should We Make Of It?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Grok 4, which has excellent benchmarks and which xAI claims is ‘the world's smartest artificial intelligence,’ is the big news.
If you set aside the constant need to say ‘No, Grok, No,’ is it a good model, sir?
My take in terms of its capabilities, which I will expand upon at great length later this week: It is a good model. Not a great model. Not the best model. Not ‘the world's smartest artificial intelligence.’ There do not seem to be any great use cases to choose it over alternatives, unless you are searching Twitter. But it is a good model.
There is a catch. There are many reasons one might not want to trust it, on a different level than the reasons not to trust models from other labs. There has been a series of epic failures and poor choices, which will be difficult to [...]
---
Outline:
(01:33) The System Prompt
(08:07) MechaHitler
(10:17) The Official Explanation of MechaHitler
(17:53) Worse Than MechaHitler
(22:22) Unintended Behavior
(26:22) Off Based
(28:37) Your Face Will Be Stuck That Way
(30:24) I Couldn't Do Solve Problem In Several Hours So It Must Be Very Hard
(38:57) Safety Third
(44:25) How Bad Are Things?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
LLMs can be deeply confusing. Thanks to a commission, today we go back to basics.
How did we get such a wide array of confusingly named and labeled models and modes in ChatGPT? What are they, and when and why would you use each of them for what purposes, and how does this relate to what is available elsewhere? How does this relate to hallucinations, sycophancy and other basic issues, and what are the basic ways of mitigating those issues?
If you already know these basics, you can and should skip this post.
This is a reference, and a guide for the new and the perplexed, until the time comes that they change everything again, presumably with GPT-5.
A Brief History of OpenAI Models and Their Names
Tech companies are notorious for being terrible at naming things. One decision that seems like the best [...]
---
Outline:
(00:51) A Brief History of OpenAI Models and Their Names
(06:05) The Models We Have Now in ChatGPT
(12:23) What About The Competition?
(12:51) Claude (Claude.ai)
(14:30) Gemini
(16:03) Grok
(16:59) Hallucinations
(19:09) Sycophancy
(20:12) Going Beyond
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Table of Contents
---
Outline:
(00:43) Language Models Offer Mundane Utility
(05:08) Language Models Don't Offer Mundane Utility
(06:57) Huh, Upgrades
(07:53) Preserve Our History
(11:18) Choose Your Fighter
(12:36) Wouldn't You Prefer A Good Game of Chess
(14:30) Fun With Media Generation
(14:40) No Grok No
(16:29) Deepfaketown and Botpocalypse Soon
(19:15) Unprompted Attention
(20:11) Overcoming Bias
(22:18) Get My Agent On The Line
(23:40) They Took Our Jobs
(27:59) Get Involved
(28:27) Introducing
(30:11) In Other AI News
(32:28) Show Me the Money
(34:59) The Explanation Is Always Transaction Costs
(37:56) Quiet Speculations
(44:23) Genesis
(46:29) The Quest for Sane Regulations
(52:06) Chip City
(52:28) Choosing The Right Regulatory Target
(01:00:42) The Week in Audio
(01:01:15) Rhetorical Innovation
(01:04:33) Aligning a Smarter Than Human Intelligence is Difficult
(01:10:10) Don't Worry We Have Human Oversight
(01:14:09) Don't Worry We Have Chain Of Thought Monitoring
(01:18:47) Sycophancy Is Hard To Fix
(01:21:43) The Lighter Side
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
It was the July 4 weekend. Grok on Twitter got some sort of upgrade.
Elon Musk: We have improved @Grok significantly.
You should notice a difference when you ask Grok questions.
Indeed we did notice big differences.
It did not go great. Then it got worse.
That does not mean low quality answers or being a bit politically biased. Nor does it mean one particular absurd quirk like we saw in Regarding South Africa, or before that the narrow instruction not to criticize particular individuals.
Here ‘got worse’ means things that involve the term ‘MechaHitler.’
Doug Borton: I did Nazi this coming.
Perhaps we should have. Three (escalating) times is enemy action.
I had very low expectations for xAI, including on these topics. But not like this.
In the wake of these events, Linda Yaccarino has stepped down this [...]
---
Outline:
(01:29) Finger On The Scale
(05:06) We Got Trouble
(07:52) Finger Somewhere Else
(09:32) Worst Of The Worst
(11:16) Fun Messing With Grok
(14:06) The Hitler Coefficient
(20:20) MechaHitler
(21:42) The Two Groks
(22:41) I'm Shocked, Shocked, Well Not Shocked
(24:05) Misaligned!
(31:39) Nothing To See Here
(33:17) He Just Tweeted It Out
(36:05) What Have We Learned?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Springtime in DC for Balsa, by Jennifer Chen
---
Outline:
(00:33) Springtime in DC for Balsa, by Jennifer Chen
(03:35) Why was no one else talking about this?
(07:27) How Likely Did We Think This Was Going to be Enacted?
(08:50) Okay, Balsa Should Do Something
(10:07) Balsa at the USTR Public Hearings
(12:34) Did Balsa... Do Anything?
(13:50) Should Balsa Continue to Do Things?
(15:26) Balsa Research is Once More 100% Focused on Jones Act Reform
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
The epic 18k word writeup on Austin's flagship Alpha School is excellent. It is long, but given the blog you’re reading now, if you have interest in such topics I’d strongly consider reading the whole thing.
One must always take such claims and reports with copious salt. But in terms of the core claims about what is happening and why it is happening, I find this mostly credible. I don’t know how far it can scale but I suspect quite far. None of this involves anything surprising, and none of it even involves much use of generative AI.
Rui Ma here gives a shorter summary and offers takeaways compatible with mine.
Table of Contents
---
Outline:
(00:47) What Is It?
(05:00) What It Isn't
(08:56) Intrinsic Versus Extrinsic Motivation
(13:30) High Versus Low Structure Learners
(14:06) I've Got a Theory
(17:54) Is This Really The True Objection?
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
Abundance and YIMBY are on the march. Things are looking good. The wins are each small, but every little bit helps. There are lots of different little things you can do. In theory you have to worry about a homeostatic model where solving some problems causes locals to double down on other barriers, but this seems to not be what we see.
There are definitely important exceptions. Los Angeles is not so interested in rebuilding from the fires and backpaddled the moment developers started to actually build 100% affordable housing because somehow that was a bad thing. New York's democratic party nominated who they nominated. Massachusetts wants to seal eviction records.
Overall, though, it's hard not to be hopeful right now. Even when we see bad policies, they are couched increasingly in the rhetoric of good goals and policies. In the long term, that leads to wins.
[...]---
Outline:
(01:11) Rent Control
(02:52) Affordable Housing
(09:19) A Vision
(09:56) Private Equity
(11:14) Home for Rent
(14:08) Making Housing Worse On Purpose So You Can Click
(16:47) Open Philanthropy Strikes Again
(18:08) The Abundance Debate
(21:24) Single Staircase Apartment Buildings
(25:18) Dublin
(25:43) Western Housing Costs
(27:22) Los Angeles
(28:16) LA Fire
(31:15) San Francisco
(34:36) California
(39:27) Oregon
(41:34) Montana
(43:58) Maine
(44:50) North Carolina
(45:10) New York City
(49:10) Massachusetts
(51:43) Texas
(54:04) Poland
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
The big AI story this week was the battle over the insane AI regulatory moratorium, which came dangerously close to passing. Ultimately, after Senator Blackburn realized her deal was no good and backed out of it, the dam broke, and ultimately the Senate voted 99-1 to strip the moratorium out of the BBB. I also covered last week's hopeful house hearing in detail, so we can remember this as a reference point.
Otherwise, plenty of other things happened but in terms of big items this was a relatively quiet week. Always enjoy such respites while they last. Next week we are told we are getting Grok 4.
Table of Contents
---
Outline:
(00:50) Language Models Offer Mundane Utility
(02:39) It Is I, Claudius, Vender of Items
(05:36) Language Models Don't Offer Mundane Utility
(06:10) GPT-4o Is An Absurd Sycophant
(06:50) Preserve Our History
(08:03) Fork In The Road
(09:43) On Your Marks
(11:06) Choose Your Fighter
(12:32) Deepfaketown and Botpocalypse Soon
(15:19) Goodhart's Law Strikes Again
(18:51) Get My Agent On The Line
(21:09) They Took Our Jobs
(22:36) Get Involved
(23:50) Introducing
(25:25) Copyright Confrontation
(27:00) Show Me the Money
(33:15) Quiet Speculations
(34:29) Minimum Viable Model
(39:14) Timelines
(41:25) Considering Chilling Out
(46:17) The Quest for Sane Regulations
(51:28) The Committee Recommends
(58:00) Chip City
(01:00:53) The Week in Audio
(01:05:33) Rhetorical Innovation
(01:15:53) Please Speak Directly Into The Microphone
(01:17:08) Gary Marcus Predicts
(01:22:10) The Vibes They Are A-Changing
(01:27:03) Misaligned!
(01:28:52) Aligning a Smarter Than Human Intelligence is Difficult
(01:32:45) The Lighter Side
The original text contained 1 footnote which was omitted from this narration.
---
First published:
Source:
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
From the publisher's feed

26,250 Listeners

2,452 Listeners

1,089 Listeners

109 Listeners

289 Listeners

90 Listeners

572 Listeners

5,556 Listeners

137 Listeners

13 Listeners

140 Listeners

145 Listeners

455 Listeners

0 Listeners

142 Listeners