
Sign up to save your podcasts
Or


The Chinese AI ecosystem has taken the AI world by storm this summer with an unrelenting pace of stellar open model releases. The flagship releases that got the most Western media coverage are the likes of Qwen 3, Kimi K2, or Zhipu GLM 4.5, but there is a long-tail of providers close behind in both quality and cadence of releases.
In this post we rank the top 19 Chinese labs by the quality and quantity of contributions to the open AI ecosystem — this is not a list of raw ability, but outputs — all the way from the top of DeepSeek to the emerging open research labs. For a more detailed coverage of all the specific models, we recommend studying our Artifacts Log series, which chronicles all of the major open model releases every month. We plan to revisit this ranking and make note of major new players, so make sure to subscribe.
At the frontier
These companies rival Western counterparts with the quality and frequency of their models.
DeepSeek
deepseek.com | 🤗 deepseek-ai | X @DeepSeek_AI
DeepSeek needs little introduction. Their V3 and R1 models, and their impact, are still likely the biggest AI stories of 2025 — open, Chinese models at the frontier of performance with permissive licenses and the exposed model chains of thought that enamored users around the world.
With all the attention following the breakthrough releases, a bit more has been said about DeepSeek in terms of operations, ideology, and business model relative to the other labs. They are very innovative technically and have not devoted extensive resources to their consumer chatbot or API hosting (as judged by higher than industry-standard performance degradation).
Over the last 18 months, DeepSeek was known for making “about one major release a month.” Since the updated releases of V3-0324 and R1-0528, many close observers have been surprised by their lack of contributions. This has let other players in the ecosystem close the gap, but in terms of impact and actual commercial usage, DeepSeek is still king.
An important aspect of DeepSeek’s strategy is their focus on improving their core models at the frontier of performance. To complement this, they have experiments using their current generation to make fundamental research innovations, such as theorem proving or math models, which ultimately get used for the next iteration of models. This is similar to how Western labs operate. First, you test a new idea as an experiment internally, then you fold it into the “main product” that most of your users see.
DeepSeekMath, for example, used DeepSeek-Coder-Base-v1.5 7B and introduced the now famous reinforcement learning algorithm Group Relative Policy Optimization (GRPO), which is one of the main drivers of R1. The exception to this (at least today) is Janus, their omni-modal series, which has not been used in their main line.
Qwen
qwenlm.ai | 🤗 Qwen | X @Alibaba_Qwen
Tongyi Qianwen, the primary AI lab within Alibaba’s cloud division, is by far and away most known for their open language model series. They have been releasing many models across a range of sizes (quite similar to Llama 1 through 3) for years. Recently, their models from Qwen 2.5 and Qwen 3 have had accelerating market share among AI research and startup development.
Qwen is closer to American Big Tech companies than to other Chinese AI labs in terms of releases: They are covering the entire stack, from VLMs to embedding models, coding models, image and video generation, and so on.They also cater to all possible customers (or rather every part of the open community) by releasing capable models of all sizes. Small dense models are important for academia to run experiments and for small/medium businesses to power their applications, so it comes to no surprise that Qwen-based models are exploding in popularity.
On top of model releases for everyone, they also focused on supporting the (Western) community, releasing MLX and GGUF versions of their models for local usage or a CLI for their coding models, which includes a generous amount of free requests.
Unlike some American companies, the core team seems to have stayed relatively small in terms of headcount, in line with other Chinese AI labs: Qwen3 has 177 contributors, whereas Llama 3 has thrice the amount, while Gemini 2.5 has over 3,000 people as part of the model.
Close competitors
These companies have recently arrived at the frontier of performance and we will see if they have the capability to consistently release great models at a pace matching Qwen or DeepSeek.
Moonshot AI (Kimi)
moonshot.cn | 🤗 moonshotai | X @Kimi_Moonshot
Moonshot AI is one of the so-called “AI tigers”, a group of hot Chinese AI startups determined by Chinese media and investors. This group consists of Baichuan, Zhipu AI, Moonshot AI, MiniMax, StepFun, and 01.AI — most of which have attracted investments by tech funds and other tech grants. For example, Alibaba is seen as a big winner in the AI space by having their own models and by being a lead investor in Moonshot, sort of like how big tech companies in the U.S. are investing in fundraising rounds for newer AI labs.
While their first models, K1 and K1.5, were closed and available on their API, they started releasing open models after the R1 release with experimental models using the Muon optimizer. Similar to DeepSeek, they focus on a single model line, with small experiments eventually feeding back into the main model. K2 is their “moonshot run,” a.k.a. yolo run, and quickly became a hit similar to R1 (see our report from the release).
Further reading on Kimi can be found on ChinaTalk.
Zhipu / Z.AI
z.ai | 🤗 zai-org | X @Zai_org
Zhipu, known in the west as Z.ai, is a startup spinoff of Tsinghua University with considerable investments by Chinese companies and VCs. Currently, they are even considering an IPO, which would make them the first AI tiger to do so.
In terms of models, they are mostly known for their recent release of GLM-4.5 and GLM-4.5V, which are all very capable for their sizes (both of which are fairly large mixture of expert models). However, they are not just releasing LLMs, but also image and video generation models, setting them apart from pure-LLM companies and labs.
Noteworthy
These companies are transitioning to open releases, have open models with inferior capabilities, or slightly different foci than the text-centric labs pushing the frontiers of intelligence.
StepFun
stepfun.ai | 🤗 stepfun-ai | X @StepFun_ai
StepFun first started as a closed model provider, but pivoted to open model releases after DeepSeek R1 shook up the industry. They are mostly focusing on multi-modal model releases, with Step3 being their flagship VLM. They also have image, audio and video generation models.
Tencent (Hunyuan)
hunyuan.tencent.com | 🤗 Tencent | X @TencentHunyuan
Hunyuan is mostly known for HunyuanVideo and Hunyuan3D. While they have released three series of different LLMs, their releases come with very strict licenses, which is unusual for Chinese companies and dampens excitement when combined with performance levels that can be found elsewhere.
RedNote (Xiaohongshu)
xiaohongshu.com | 🤗 rednote-hilab
The Chinese version of Instagram, RedNote, recently joined the ranks of Chinese companies releasing open models. Especially their capable character recognition / OCR model surprised many (see our coverage). Similar to Xiaomi and Baidu, it remains to be seen what their overall open strategy will be in the near and distant future and they have not competed in the large, frontier model space.
MiniMax
minimaxi.com | 🤗 MiniMaxAI | X @MiniMax__AI
MiniMax is another of the AI tigers and also started as a closed company. After the release of R1, they changed their strategy and released the weights of Minimax-Text-01, following up with reasoning models building upon it. The unique selling point of these models are the 1M context window achieved with hybrid attention.
These text models are not the only thing they are focusing on — they also have image and video generation models, but those remain closed and only available on their API. They are also promoting their consumer platform heavily as they eye an IPO.
OpenGVLab / InternLM
internlm.intern-ai.org.cn | 🤗 InternLM | X @opengvlab
InternLM & OpenGVLab have deep ties to the Shanghai AI Laboratory, with InternLM focusing on the language models, while OpenGVLab releases vision models. While they release a range of models such as S1 or InternLM-Math, the orgs are mostly known for the strong InternVL series. While the first versions mostly used their own InternLM pretrained models, later releases (such as InternVL3) rely on Qwen as the language backend.
Skywork
skywork.ai | 🤗 Skywork | X @Skywork_AI
The Singaporean Skywork first started out as an online karaoke company (yes, really) before they pivoted to AI and being a competitor to Manus, with their platform focusing on agents for work-related tasks, such as slide generation.
Their LLM journey started with them releasing their own pretrained dense and MoE models. However, they stopped pre-training their own models and instead started to fine-tune existing models: Their OR1 reasoning model builds on top of DeepSeek-R1-Distill-Qwen-32B, R1V3 uses InternVL3 (which itself uses Qwen2.5 as its LLM backend).
Aside from LLMs, they have a wide range of other models, from world models, image and video generation models, and reward models. Similar to their LLMs, they mostly build on top of other models. Unlike many labs, Skywork has released some datasets with their models, such as preference and reasoning training data.
On the rise
These companies are either just getting their toes wet with open models or operating as more of academic research organizations than labs pushing the performance of models.
ByteDance Seed
seed.bytedance.com | 🤗 ByteDance-Seed
Seed is the R&D arm of ByteDance and eerily similar to Meta’s FAIR division: Diverse models with interesting research, with their papers garnering a ton of attention in the community. However, it remains to be seen whether they shoot for a Llama-style model release or continue to release research artifacts.
Here are some recent papers:
* Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference
* Seed-Prover: Deep and Broad Reasoning for Automated Theorem Proving
* Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
* Seedance 1.0: Exploring the Boundaries of Video Generation Models
* SeedEdit 3.0: Fast and High-Quality Generative Image Editing
* Seed1.5‑VL Technical Report
* Mogao: An Omni Foundation Model for Interleaved Multi‑Modal Generation
* Seed1.5‑Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
* VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
* Seed LiveInterpret 2.0: End‑to‑end Simultaneous Speech‑to‑speech Translation with Your Voice
OpenBMB
openbmb.ai | 🤗 openbmb | X @OpenBMB
OpenBMB is an open-source community (comparable to BigScience) from Tsinghua University NLP Lab (the very same university where Zhipu was spun off from) with support from the Beijing Academy of Artificial Intelligence (BAAI) and ModelBest.
They are mostly focusing on small multi-modal models for the edge, such as MiniCPM-V-4. However, the license is rather restrictive, which is surprising given the community-driven origins of the group. Aside from model releases, they also release frameworks and specialized kernels to make sure their models run on low-end hardware.
Xiaomi (MiMo)
mi.com | 🤗 XiaomiMiMo
Xiaomi started releasing a bunch of small, capable models, ranging from LLMs to VLMs and audio models. Xiaomi updating the models quickly after an initial launch and releasing multiple variants of the models show that it is not a one-off foray into open models. However, it remains to be seen whether those are mostly research artifacts or whether they are serious about potentially pushing the frontier or competing for adoption.
Baidu (ERNIE)
yiyan.baidu.com | 🤗 baidu | X @Baidu_Inc
Baidu, one of the original names in the Chinese AI space, has only released the weights of ERNIE 4.5. It remains to be seen whether they continue to release weights of newer releases as well.
Honorable Mentions
The rest of the labs that we are watching.
Multimodal Art Projection
m-a-p.ai | 🤗 m-a-p
An open research community, releasing all kinds of models (including a truly open 7B language model with data, etc.). Now, they’re mostly known for the music generation model YuE.
Alibaba International Digital Commerce Group
aidc-ai.com | 🤗 AIDC-AI
Another R&D arm of Alibaba, mostly releasing niche models building upon Qwen.
Beijing Academy of Artificial Intelligence (BAAI)
baai.ac.cn | 🤗 BAAI | X @BAAIBeijing
As a university, the Beijing Academy of Artificial Intelligence has a high diversity of projects. They are mostly known for BGE, which are capable embedding models.
inclusionAI
🤗 inclusionAI | X @InclusionAI666
The open weight arm from the Ant Group (an affiliate of Alibaba handling mobile payments and some financial industries), responsible for Ling Lite, a series of LLMs.
Pangu (Huawei)
huaweicloud.com | X @HuaweiCloud1
Huawei is working on AI accelerators to threaten the market share of Nvidia GPUs, which are often targeted by regulations, both from the US and China. Their model releases are mostly to show what’s possible with their cards, but not without drama accusing them of upcycling Qwen models and not stating it. We would expect them to continue to release more models in the near future.
Dwarkesh Patel’s now well-read post on why he is extending his AI timelines focuses on the idea of continual learning. If you ask me, what we have already is AGI, so the core question is: Is continual learning a bottleneck on AI progress?
In this post, I argue that continual learning as he describes it actually doesn’t matter for the trajectory of AI progress that we are on. Continual learning will eventually be solved, but in the sort of way that a new type of AI will emerge from it, rather than continuing to refine what it means to host ever more powerful LLM-based systems.
Continual learning is the ultimate algorithmic nerd snipe for AI researchers, when in reality all we need to do is keep scaling systems and we’ll get something indistinguishable from how humans do it, for free.
To start, here’s the core of the Dwarkesh piece as a refresher for what he means by continual learning.
Sometimes people say that even if all AI progress totally stopped, the systems of today would still be far more economically transformative than the internet. I disagree. I think the LLMs of today are magical. But the reason that the Fortune 500 aren’t using them to transform their workflows isn’t because the management is too stodgy. Rather, I think it’s genuinely hard to get normal humanlike labor out of LLMs. And this has to do with some fundamental capabilities these models lack.
I like to think I’m “AI forward” here at the Dwarkesh Podcast. I’ve probably spent over a hundred hours trying to build little LLM tools for my post production setup. And the experience of trying to get them to be useful has extended my timelines. I’ll try to get the LLMs to rewrite autogenerated transcripts for readability the way a human would. Or I’ll try to get them to identify clips from the transcript to tweet out. Sometimes I’ll try to get them to co-write an essay with me, passage by passage. These are simple, self contained, short horizon, language in-language out tasks - the kinds of assignments that should be dead center in the LLMs’ repertoire. And they're 5/10 at them. Don’t get me wrong, that’s impressive.
But the fundamental problem is that LLMs don’t get better over time the way a human would. The lack of continual learning is a huge huge problem. The LLM baseline at many tasks might be higher than an average human's. But there’s no way to give a model high level feedback. You’re stuck with the abilities you get out of the box. You can keep messing around with the system prompt. In practice this just doesn’t produce anything even close to the kind of learning and improvement that human employees experience.
The core issue I have with this argument is the dream of making the LLMs we’re building today look more like humans. In many ways I’m surprised that Dwarkesh and other very AGI-focused AI researchers or commentators believe this — it’s the same root argument that AI critics use when they say AI models don’t reason. The goal to make AI more human is constraining the technological progress to a potentially impossible degree.
Human intelligence has long been the inspiration for AI, but we have long surpassed it being the mirror we look to for inspiration. Now the industry is all in on the expensive path to make the best language models it possibly can. We’re no longer trying to build the bird, we’re trying to transition the Wright Brothers’ invention into the 737 in the shortest time frame possible.
To put it succinctly. My argument very much rhymes with some of my past writing.
Do language models reason like humans? No. Do language models reason? Yes.
Will language model systems continually learn like humans? No.Will language model systems continually learn? Of course.
Interconnects is a reader-supported publication. Consider becoming a subscriber.
Dwarkesh writes “Rather, I think it’s genuinely hard to get normal humanlike labor out of LLMs.” This is because we’re still early on the buildout of the technology. Human labor takes an immense amount of context and quick thinking, both of which we’re starting to unlock with our language models. On top of this, human labor may not be what we want to create — we want to augment it.
Using LLMs as drop in replacements for humans is not a requirement for AGI nor is what Dwarkesh describes a fundamental limitation on AI progress. Francois Chollet cleverly poked at this weakness in his recent conversation with Dwarkesh at an ARC-AGI event:
Well, how do you define the difference between the ability to adapt to a new task and learning on the fly? It's, it sounds like the same thing to me.
Language models can already pick up subtle context extremely fast. ChatGPT’s memory feature has gotten far better for me. When we’re using the far more powerful models we can expect in the next 18 months this’ll already start to appear magical. Language models are extremely apt at inferring context even without us giving it to them. Soon we’ll be unlocking that subtle connection engine by providing immense, explicit context.
I don’t know of anyone who has actually thoroughly digitized all the relevant context of their job and formatted it in a way that is easily readable by an LLM. GPT-5 Pro estimates that all of the writing on Interconnects would be only 500K tokens. That would fit into an existing LLM with no extra system, but I’ve never tried it.
The problem that Dwarkesh is facing is that we’re still using LLMs primarily in a single generation manner, which got far better with the introduction of reasoning models, but the economically useful way to use current tools in more complex intellectual domains will require a deep-research style approach over all of your recent work interactions. No one is giving language models that kind of context. None of the tools we use are set up properly to accumulate this type of context.
I expect this to change rapidly. ChatGPT, Claude, and the likes are all adding memory features across chats and countless connectors to other pieces of information in your professional life. These memory features will be omnimodal and essential to extracting the type of value Dwarkesh wants. Without them, I agree language models in their current form are hopeless at solving continual learning.
This is what I would expect the rumored $2000/month ChatGPT level subscriptions to work with. Each of these bespoke tasks needs to absorb a ton of context and reasoning tokens in order to make a directionally right output. If someone built the Claude Code equivalent for my Substack, with every post tagged by topic and performance metrics, I bet the AI could easily make useful suggestions on how to format my content.
Continual learning in how Dwarkesh presents it is a systems problem rather than a learning problem. I expect better context management over my information ecosystem to exist in 2026, but more work to be needed for the AI companies to know how best to reference it and unlock in-context learning that feels like rapid adaptation. Call that 2027.
The models that have been released in 2025 will make this far more tractable in the near future. Reasoning models have made in-context learning far more powerful, resulting in rapid progress on held-out and complex domains such as ARC-AGI. These models also have come with massive improvements in context length. Claude and Gemini have 1M+ token context lengths and GPT-5’s is at 400K — they’re all growing steadily. What is important with the context length numbers is that evaluations are showing that these are meaningful improvements that the models can leverage intelligently.
With these reasoning models and smart retrieval of context, the systems we are building will look indistinguishable from continual learning. This will definitely be multiple LLMs working together and will operate very differently than the first versions of ChatGPT we were given (and often still use today).
The path to continual learning is more context and more horsepower. This is directly in line with the direction AI investment is going. This doesn’t feel like a bottleneck, rather another product problem that we are going to solve. This sort of continual learning may not enable the type of raw intelligence and autonomy that many vocal leaders in AI describe as “superintelligence.”
Training models to be smarter on even more complex tasks — e.g. novel biological research — requires mastering agentic behaviors that need to be learned from scratch, as discussed in my post on “What comes next with RL”. There’s no internet scale pretraining data for such agentic tasks. My point is that not all jobs that require continual learning will require the frontiers of intelligence. I’m excited to write blog posts with the bliss of my ChatGPT 6 co-editor.
This technology coming soon will not be without its challenges. My first reaction to the continual learning post was more in line with “society isn’t ready for this” rather than commentary on its feasibility. I’ll repeat my warning:
For a long time I’ve written that AI models have a higher risk potential in terms of social outcomes because the modalities they interact with us in are far more personal… As AI is going to be so powerful as a standalone entity, breaking some of the symbiotic links will be good for adding friction that makes the technology easier to steer towards good outcomes. In short, be wary of wishing for end-to-end (reinforcement) learning when you’re part of the environment.2 It’s a destiny to dystopia.
What we have today is a form of AGI and it’ll soon get much better with better context and memory. The industrialization of language models is giving us incredible improvements across a wide swath of use-cases. These will blow past many basic primitives of intelligence in humans that have motivated AI for decades. First was models reasoning, then will come systems with continual learning. This is exactly what most AI companies are actually building — regardless of what their superintelligence messaging is.
Comments are open on this post, please continue the debate!
If you want a video version of this, check out the last 20 minutes of the livestream reaction (edit, fixed link) I did with Will Brown of Prime Intellect and Swyx of Smol AI & Latent Space.
GPT-5 was set up to fail on some of the narratives it was expected to satisfy. The two central themes it had to decide between were the AGI (or superintelligence) narrative that Sam Altman & co. have been using to fundraise and the fact that ChatGPT is one of the fastest-growing consumer technologies of all time.
To fulfill both, GPT-5 needed to be AGI while also being cheap enough to serve as the most-used AI system in the world. Business and technological realities made it inevitable that GPT-5’s primary impact would be to solidify OpenAI’s market position, even if it raises a lot of eyebrows for the long-term trajectory of AI.
The reactions online capture this as well. The OpenAI live streams have historically catered to AI insiders, but the product speaks entirely to a different audience. The people discussing this release on Twitter will be disappointed in a first reaction, but 99% of people using ChatGPT are going to be so happy about the upgrade. Confusingly enough, this includes many of the critics. GPT-5 is a good AI system. It’s right in line with best-in-class across pretty much every evaluation, while being cheap enough to serve the whole world.
OpenAI is largely fixing its product offering with an announcement that was hyped to be one of the biggest AI news cycles of the year. AI news being loud is defined by narratives being different more-so than technology being better. OpenAI releasing an open model again will likely be pinpointed as just as important a day for the arc of AI as the GPT-5 release. In many ways GPT-5 was set up to fail and that is very off-putting for those expecting maximum AI progress in the near term.
I’m not going to dwell on it, but oh boy, that was a messy release. GPT-5 being announced and rolled out like this is very odd. Countless plots were mislabeled, live demos had bugs, and the early rollout is doing some weird stuff. This reinforces how OpenAI was torn about the release and backed into a corner with their messaging. They knew they needed to improve the experience with strong competition in the industry, but releasing GPT-5 needed to make a splash after how long they’ve waited (and already parked the GPT 4.5 name).
The core question we track in this post is: What does it mean for the next 6-18 months of AI progress if GPT-5 is just as good as all the best models out there, e.g., Claude Sonnet for coding or o3 for search, funneled into one, super cheap package?
If AGI was a real goal, the main factor on progress would be raw performance. GPT-5 shows that AI is on a somewhat more traditional technological path, where there isn’t one key factor, it is a mix of performance, price, product, and everything in between.
Interconnects is a reader-supported publication. Consider becoming a subscriber.
GPT-5’s performance
There are a few places that we can see that GPT-5 represents a solid step on the performance trend line, but nothing like a step change. First, on LMArena, GPT-5 is fantastic, sweeping the board to #1 on all categories. The last model to claim #1 in pretty much every category was Gemini 2.5 Pro — and that was the biggest step change in Elo since GPT-4 Turbo skyrocketed past the first Claude.
Second, GPT-5 is the top model on the ArtificialAnalysis composite benchmark.
These two, LMArena & ArtificialAnalysis, represent two coarse evaluations — community vibes and raw benchmarks. Both of these can be gamed, but are still correlated with real-world use. You can also see in OpenAI’s shared results how much the smaller versions improve on the likes of GPT-4.1 mini and o4-mini.
In many ways, the march of progress on evals has felt slowed for a while because model releases are so frequent and each individual step is smaller. Lots of small steps make for big change. The overall trend line is still very positive, and multiple companies are filling in the shape of it.
My post on “what comes next” from earlier this summer all but called this type of release, where the numbers aren’t shocking but the real world use cases are great, becoming more common.
This is a different path for the industry and will take a different form of messaging than we’re used to. More releases are going to look like Anthropic’s Claude 4, where the benchmark gains are minor and the real world gains are a big step. There are plenty of more implications for policy, evaluation, and transparency that come with this. It is going to take much more nuance to understand if the pace of progress is continuing, especially as critics of AI are going to seize the opportunity of evaluations flatlining to say that AI is no longer working.
To say it succinctly: Abilities will develop more slowly than products.
The product overhang is being extended with each release. We’re still building untapped value with AI models and systems faster than we’re capturing it.
Another way to see this incremental push out in models or systems is through OpenAI’s update to the famous METR plot of time to completion for humans of various tasks AI systems can solve 50% of the time. GPT-5 is leading, but also just in line with trends.
All of this is to say comprehensively that AI progress is very alive and well, as long as you don’t subscribe to the exponential takeoff in ability. Those arguments are very strained by this GPT-5 release.
Yes, AI progress on intelligence and “raw ability” is certainly going to continue at a solid pace for a long time, but how will this translate into recursive self-improvement?
GPT-5’s details
If you’re reading closely, you may have noticed that this post uses the word system instead of model. All of the leading chat systems have been adding more components onto them like safety checkers and so on, but this is the first one to use different architectures and weights for the primary generation of content across similar queries. GPT-5 is the first in what is to come, mostly to better balance cost and give better user experiences. From the system card:
GPT‑5 is a unified system with a smart and fast model that answers most questions, a deeper reasoning model for harder problems, and a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent (for example, if you say “think hard about this” in the prompt). The router is continuously trained on real signals, including when users switch models, preference rates for responses, and measured correctness, improving over time.
Along with this, they shipped many product improvements, such as how the model has a 400K context window in the API with great performance, reduced hallucinations, and new personalities.
Primarily, I worry as a power user about the router. I sense that for now I’ll default to GPT-5 Thinking, and sometimes upgrade to Pro mode, while downgrading to standard GPT-5 only for benign queries (depending on its search behavior — if it is search-heavy like o3 without thinking, then it should still work well).
Thankfully, the thinking mode has a “get an early answer” button, so I don’t see any reason to start elsewhere. If I need an answer fast, I’ll get one. If not, I want the best responses possible.
As for prices, here’s a comparison. GPT-5’s top-level model is cheaper than Claude Sonnet and far better than any OpenAI model has been before at coding — one of the core details of this release. Matching Gemini Pro’s pricing when considering Google’s infrastructure advantage is a substantial accomplishment.
* OpenAI — GPT-5 (API sizes)
* GPT-5: input $1.25, output $10.00. (OpenAI)
* GPT-5 mini: input $0.25, output $2.00. (OpenAI)
* GPT-5 nano: input $0.05, output $0.40. (OpenAI)
* OpenAI — o3 (reasoning)
* o3: input $2.00, output $8.00. (OpenAI Platform)
* o3-mini: input $1.10, output $4.40. (cached input $0.55) (OpenAI Platform)
* Anthropic — Claude 4 family
* Claude Sonnet 4: input $3.00, output $15.00. (Anthropic)
* Claude Opus 4.1: input $15.00, output $75.00. (Anthropic)
* Google — Gemini 2.5
* Gemini 2.5 Pro: input $1.25 (≤200k prompt) / $2.50 (>200k); output $10.00 (≤200k) / $15.00 (>200k). (Google AI for Developers)
* Gemini 2.5 Flash: input $0.30 (text/image/video) or $1.00 (audio); output $2.50 (includes thinking tokens). (Google AI for Developers)
* Gemini 2.5 Flash-Lite: input $0.10 (text/image/video) or $0.30 (audio); output $0.40. (Google AI for Developers)
Cheaper, thinking models that work well in applications are far more useful than scaling (as GPT-4.5 has shown us).
GPT-5’s impact
It seems like most people in all walks of life are going to love this model — from AI researchers all the way to people who are learning of ChatGPT for the first time today. This is very in line with my expectations for how AI will proceed, as a long, steady march of progress.
The fact that the models are getting way cheaper rather than way more expensive definitely signals that we cannot just brute-force scale our way to much stronger systems. Scaling helps, but it is now one of many considerations, and all the laboratories are showing us that much bigger models have diminishing returns in value to customers. At the same time, models being cheaper could be just what we need for Jevons paradox to kick in and provide another boost in AI adoption.
Many people will claim that the GPT-5 release was a flop and the bubble will pop for AI. This is downstream of the industry generally making totally unrealistic promises. As someone whose core through-line when covering frontier models is tracking the pace of progress, I translate this as “AI capabilities on benchmarks will proceed a bit more slowly, but we aren’t reaching any clear walls in performance.” The AI performance hills we’re climbing up as an industry do put up some more resistance as the obvious low hanging fruit is gone, but we have the tools to overcome it consistently for the next 6 to 18 months.
For companies that have been fundraising on promises of AGI, such as Anthropic and OpenAI, closing the next rounds could be harder. Of course, this depends on whether the messaging of the rounds was a key part of the fundraising.
This fundraising inspires capital expenditures across the industry, e.g. TSMC developing the next node for NVIDIA to build new chips, and so on. The AGI narrative and the fundraising it has enabled have been good for the U.S. in terms of building out valuable, raw infrastructure.
This could be the beginning of the money train slowing down, but that’s very different from a derailment and a stock market crash. As raw infrastructure spend slows, there will be even more pressure to deliver valuable products to users. A key trend for 2025 has been many of those appearing — Deep Research and Claude Code being the paradigms that everyone has copied.
GPT-5 makes these applications better and makes it easier and cheaper for the next viral AI products to hit the market. I’m still excited for what is to come.
But first, I’m going to sign off and go play with GPT-5. It’s a good day to build something for the fun of it. As I use it more, I’ll have more to say.
Extra GPT-5 links
For more specifics on the model from people who got early access, I recommend Tyler Cowen, Every.to, or Simon Willison (or Swyx soon, on Latent.Space).
Livestream link: https://openai.com/gpt-5/ Research blog post: https://openai.com/index/introducing-gpt-5/ Developer blog post: https://openai.com/index/introducing-gpt-5-for-developers Enterprise blog post: https://openai.com/index/gpt-5-new-era-of-work GPT-5 landing page: https://openai.com/gpt-5/ System Card: https://openai.com/index/gpt-5-system-card/ Coding examples: https://openai.github.io/gpt-5-coding-examples/What would you say if you could talk to a future OpenAI model https://progress.openai.com/
Finally, I’ll plug again the video I did with Will Brown and Swyx:
Send me the most interesting things you find on GPT-5!
OpenAI released two open-weight, text-only reasoning models today, both mixture of experts (MoE) sized to run efficiently on a range of hardware from consumer GPUs to the cloud. These models have the Apache 2.0 license, so they’re available for distillation into other reasoning models, deployment into commercial products, and are free of downstream restrictions. These two models, the smaller gpt-oss-20B with 3.6B active parameters and 21B total and the larger gpt-oss-120B with 5.1B active parameters, follow the trends we’ve seen with the other leading open models in architecture choices.
Where this release shines is in the dramatic change in open model performance and strategy that comes with the leading name in AI releasing an open model that undercuts some of their own API products.
We’ll get to the technical details on the model later, but the main point of this post is how much OpenAI has changed by releasing their first open language model since GPT-2. The larger 120B model “achieves near-parity with OpenAI o4 mini on core reasoning benchmarks” and is a major moment for the ecosystem:
* OpenAI has released an open model at the frontier of current open model performance — highlighting how major concerns over open models that OpenAI leadership mentioned in 2023 were overblown. The marginal risks of open models have been shown to not be as extreme as many people thought (at least for text only — multimodal is far riskier). Once other organizations, particularly Meta and China showed OpenAI that there was no risk here, the path was opened to release a model.
* OpenAI has revealed far more of their technical stack than any release to date. This blog post has light details on many things in the model, but community tinkering will begin to better understand what is going on here. This includes basic things like our first time seeing a raw chain of thought (CoT) for an OpenAI reasoning model, but also more interesting things like how this model is trained to use tools in the CoT like their o3 model. Other details include researchers being able to play with OpenAI’s instruction hierarchy in raw weights (where pieces of it are untouchable in the API), a new “harmony” prompt format, the same “reasoning efforts” of low, medium & high from the API, a huge proof of concept on how far basic, community standard architectures with MoEs can be pushed, and other small details for the AI community to unpack.
* OpenAI has initiated a scorched earth policy on the API market, undercutting their own offerings and unleashing an extremely strong, trusted model brand with a permissive license. While adoption of any open model is much slower than an API due to testing, additional configuration, etc., this is set up to go about as fast as it can. Any API model that competes with current models like OpenAI o4 mini, Claude Haiku, Gemini Flash, DeepSeek R1 etc. are all going to have to compete with this model. OpenAI’s o4 mini model is currently served at $1.1 per million input tokens and $4.4 per million output. Serving this open model will likely cost at least 10x less. There are many potential strategic reasons for this, all of which paint OpenAI as having a clearer vision of what makes it valuable. What OpenAI hasn’t touched with this model is interesting too — “For those seeking multimodal support, built-in tools, and seamless integration with our platform, models available through our API platform remain the best option.” These are dropped for reasons above, and “headaches” discussed later in the post.
Together, these paint a much clearer vision by OpenAI on how they’ll control the AI ecosystem. The top potential reasons on my mind are:
* OpenAI could be trying to make all API models potentially obsolete on cost ahead of the GPT-5 release, which they hope to capture the top end of the market on. Or,
* OpenAI could be realizing that models are no longer their differentiation, as ChatGPT users continue to steadily climb — and they’ll soon pass 1 billion weekly actives.
There are plenty of other reasons, such as the politics alluded to at the end of the blog post, but OpenAI tends to only act when it serves them directly — they’ve always been a focused company on their goals.
There’s also a long list of head scratchers or in-between the lines points that illuminate OpenAI’s strategy a bit more. OpenAI of course didn’t release training data, code, or a technical report, as expected. OpenAI is trying to make a big splash with the name that captures more of the enterprise market, but in doing so takes some collateral damage in the research and true “open source” AI communities. These future questions include:
* The naming is bad — a mixture of cringe, confusion-inducing, and still useful for their marketing goals. For anyone following open-source AI for a long time it won’t be new that a major company is blurring the association of the term open-source with the community accepted definitions. I understand why OpenAI did this, but the naming conflict further enforces that the true open source AI community isn’t the target of this release — it’s people that want to try an “open source AI model” for their business, and OpenAI has made the target too big to miss for enterprises.
* OpenAI did not release the base models. Anyone following the space would’ve expected this, but it matters substantially for researchers. These two sparse, low numerical precision MoE models won’t be easy for researchers to use. The best model for researchers and tinkerers are dense, base models from 1 to 7 billion parameters. These are much “longer term” artifacts in the open community that will still be using almost only Qwen.
I need to take a second before the “unknowns” section and comment on the architecture. These models are reinforcing trends we’re seeing in modeling across the industry. Recent frontier open models are all very sparse MoEs inspired by the DeepSeek architecture. DeepSeek V3 had 37B active and 671B total parameters. Kimi K2 had 32B active and 1T total parameters. With 5B active and 121B total, the sparsity factor fits right in with normal. Sparsity in MoEs is totally king right now. The smaller gpt-oss is a bit less sparse than Qwen’s 3B active, 30B total smaller MoE, but expect the sparsity of these models to continue to increase.
Some things we need more testing to know the impact of include:
* The model has been quantized for release to MXFP4 (4 bit floating point). It’s not clear exactly who will be impacted here, but this could make it benefit people most with the newest hardware, cause minor issues across Torch/Cuda versions, or even make some of the behaviors weird relative to the trained version internal to OpenAI. This could also be a plus, depending on performance, as the bigger model is quantized to 4 bit precision to enable it to be run on GPUs with 80GB of memory, such as the A/H100 line from NVIDIA.
* Safety measures have been taken to change how finetunable the model is. With, or soon after, this release OpenAI is releasing a research paper on new methods to make it so you can’t “finetune the safety away” from a released instruct model. This is a very long-standing issue that people have concerns with over releasing open models. The main question here is if the models OpenAI releases are still able to be finetuned or not for productive use-cases. OpenAI claims they can be in their blog post, but this will be left up to the community to decide. Is finetuning the safety away actually a feature of an easy to use model?For example, Gemma has been tougher for people to finetune historically because it uses a different attention implementation and has a different parameter space from being distilled. Open finetuning stacks are still tuned for Llama and Qwen — this takes a long time to change.Many people will take the “we made it impossible to un-censor this model” as a challenge, which will be interesting to follow in the jailbreaking research community. There is a substantial market for modifiable models.
* The model was trained to expect tools, but open model tool use is a mess. One of the biggest problems I worry about in designing an OLMo model with native o3-style tool use is that I need to make it seamless for users to use the same tools from training time at inference time. An early tester in my network mentioned that the model would hallucinate tool calls from training (sort of like what was mentioned around o3’s full release). I don’t expect this to be an unsolvable issue, but it could slow adoption. It could also allow people to reverse engineer the tools that OpenAI uses during training, we’ll see!
* We need to re-benchmark the model on open infrastructure. OpenAI did a good job for this release integrating it everywhere, but we need to confirm that the community can easily replicate their evaluation scores. Evaluation at closed labs has increasingly become bespoke to suit their internal needs, which is a logical decision, but this comes at a cost of friction when an open model is released. This is me saying loud and clear that this isn’t a model performance review in a nuanced sense, but a summary of the importance of OpenAI’s approach (and where the opportunity is for the rest of us). Not all good models are easy to use. Some models benchmark well and are useful — e.g. Qwen. Some models benchmark well and are forgotten. Regardless of scores, I expect this to be a useful model.
Overall, I would give OpenAI a very strong grade on their first open release in a while — they definitely listened to the feedback given by the community. The path to earning goodwill with the open community, especially with researchers, is to embrace more risk in making models that are easier to modify (and potentially even more revealing), such as the base models for these checkpoints.
Open models from the U.S. labs were in such a dire spot that we need any step back in the right direction. As the rollout of the model begins and we have more understanding of it, we’ll include more updates on Interconnects, such as in the next Artifacts Log issue.
Interconnects is a reader-supported publication. Consider becoming a subscriber.
So, OpenAI is the new open champion, right? There’s no more risk vis-a-vis China? We don’t need Llama anymore? Not quite, let me explain.
OpenAI, ATOM, and national champions
It’s a phenomenal step for the open ecosystem, especially for the West and its allies, that the most known brand in the AI space has returned to openly releasing models. This is momentum and could be the start of the turning point of adoption and impact of open models relative to China.
The open ecosystem moves fast in some ways and slow in others. Many workflows and expertise is now built on Qwen models due to their frequent, accessible releases. Some of these will try OpenAI the next time they want to make a change, but it’s far from the fact that everyone will immediately switch to OpenAI’s model now that it’s out.
To me, OpenAI dropping a strong model has switched the second derivative on the open model scales. The U.S. and its allies will no longer be falling further and further behind, which was the main story of 2025, but we need to build on this momentum if we want to have competitive open models for all use cases in the order of months rather than years.
There’s a lot of uncertainty in the incentives for open models. Some of the best China analysts I know share how China is sensing that releasing open models is a successful strategy for them and are doubling down. This is a very reasonable take. The retort is that if we use it as a weakness of the American ecosystem that it is so reliant on Meta’s Llamas, or now GPT OSS, the same could happen for Qwen. So then, what happens if Alibaba decides Qwen’s stellar releases no longer serve them?
In this case, there would be a large opportunity in the series of small models from 1 to 70B parameters, but there’s so much competition from China at the larger scales. These are currently the big mixture of experts (MoE) models like DeepSeek V3/R1, Z.ai’s / Zhipu’s GLM 4.5, Kimi K2, and so on. China has more models that are close to this performance level, such as MiniMax or Tencent.
All of these companies have uncertainty, but there’s a strength in numbers that reinforces standard practice and sets standards. Releasing strong, large, open models is now the standard in China. We’re back in the precarious period of establishing standards for American companies, who are exposed to the legal risk of not being able to un-release models with many open lawsuits, such as in areas like copyright.
These two sides of the open ecosystem are at very different stages and need very different actions. In many ways, we shared The ATOM Project when we did because we could tell this was a local (and hopefully global) minimum in terms of the distance between Western contributions to the open science of AI compared to any point in the recent past and near future.
OpenAI’s release is a step in the right direction, but it is still a precarious position. Many people make noise about creating open models, from the AI Action Plan to venture capitalists and academics. What all of these parties have in common is that its not their number one goal. The goal of The ATOM Project is to give an outlet for people like myself that want to make this project their number one priority.
This is why we need to keep nurturing entrants into the open model space that are releasing their best models there. It is what made the early versions of Llama great, and is what will be the defining factor of the outputs of ATOM. Models that are designed from first principles to be modifiable, interpretable, and extendable is what will enable a new decade of AI research to be born. This needs base models, training details, convenient sizes, and other little details that are missing from many recent open model releases, including OpenAI’s.
I’m very excited to share a substantial project on invigorating investment in open language models and AI research in the U.S. The ATOM (American Truly Open Models) Project is the mature evolution of my original “American DeepSeek Project” and I hope it can help be a turning point in the current trajectory of losing open model relevance vis-a-vis China, and even the rest of the world.
I’ve included the full text below, but I encourage you to visit the website for the full version with added visuals, data, and a place to sign your support. This is a community movement, rather than me fundraising, starting an organization, or anything like that
If you can help get the word out and or sign your support, I’d greatly appreciate it.
(Or watch a 5 minute overview on YouTube)
The ATOM Project: Towards fully open models for US research & industry
Reinvigorating AI research in the U.S. by building leading, open models at home
America's AI leadership was built by being the global hub and leading producer of open AI research, research which led directly to innovations like the Transformer architecture, ChatGPT, and the latest innovations in reasoning models and agents. America is poised to lose this leadership to China, in a period of geopolitical uncertainty and rising tensions between these two nations. America's best AI models have become more closed and restricted, while Chinese models have become more open, capturing substantial market share from businesses and researchers in the U.S. and abroad.
Open language models are becoming the foundation of AI research and the most important tool in securing this leadership. America has lost its lead in open models – both in performance and adoption – and is on pace to fall further behind. The United States must lead AI research globally, and we must invest in making the tools our researchers need to do their job here in America: a suite of leading, open foundation models that can re-establish the strength of the research ecosystem.
Recommendation: To regain global leadership in open source AI, America needs to maintain at least one lab focused on training open models with 10,000+ leading-edge GPUs. The PRC currently has at least five labs producing and releasing open models at or beyond the capabilities of the best U.S. open model. Regaining open source leadership is necessary to drive research into fundamental AI advances, to maximize U.S. AI market share, and to secure the U.S. AI stack.
Overview
Open language model weights and data are the core currency of recent AI research – these are the artifacts that people use to come up with new architectures, training paradigms, or tools that will lead to the next paradigms in AI to rival The Transformer or Inference-time Scaling. These research advances provide continued progress on existing products or form the basis for new technology companies. At the same time, open language models create potential for a broader suite of AI offerings by allowing anyone to build and modify AI how they see fit, without their data being sent through the cloud to a few, closed model providers.
Open language models are crucial for long-term competition within American industry. Today, substantial innovation is happening inside of large, closed AI laboratories, but these groups can only cover so many of the potential ideas. These companies spend the vast majority of their resources focusing on the next model they need to train, where the broader, open research community focuses on innovations that’ll be transformative in 2, 5, 10, or more years. The most progress in building useful, intelligent AI systems will come when the most people can participate in improving today's state-of-the-art, rather than the select few at certain companies.
The open AI ecosystem (regarding the models, not to be confused with the company OpenAI) has historically been defined by many parties participating. The United States emerged as a hub of the deep learning revolution via close collaboration between leading technology companies and academic institutions. Following ChatGPT, there have been countless contributions from around the globe. This distribution of impact on research has been collapsing towards clear Chinese leadership due to their commitment to open innovation, while a large proportion of leading scientists working in the United States have joined closed research organizations.
The playbook that led Google to invent and share the Transformer – the defining language model architecture of which all leading models such as ChatGPT, Gemini, or Claude are derived from – is now the standard mode of operation for Chinese companies, but it is increasingly neglected by American companies.
The impact of China’s models and research are growing because the institutions focused on open models have access to substantial compute resources for training – e.g. some have formed a close relationship between leading AI training laboratories and academic institutions. Until the United States and its partners directly invest in training more, higher performance open models and sharing the processes to do so, its pace of progress in AI research will lag behind.
To train open models at the frontier of performance, a developer currently needs a high concentration of capital and talent. We estimate that to lead in open model development, the United States needs to invest in multiple clusters of 10,000+ H100 level GPUs to create an ecosystem of fully open language models that are designed to enable a resurgence in Western AI research. Stacking large investments such as this into a few focused efforts will help them to learn from each other and make progress across a range of challenges quickly and robustly. Splitting such an investment in AI training into smaller, widespread projects will not be sufficient to build leading models due to a lack of compute concentration. Along the way we need to build models of various sizes that can enable applications of AI at every scale from local or edge devices all the way to high performance cloud computing.
Open models as the engine for AI research and development
America's AI leadership was built by tens of thousands of our best and brightest students, academics and researchers. This process occurred over decades, but it is faltering at a crucial transition point to the new, language modeling era of AI research. Since the release of ChatGPT, open language models and computational resources are the most important table stakes for doing relevant and impactful research. High-quality open models and their subsequent technical reports quickly accrue thousands of citations and accolades such as best paper awards and the focus of large swaths of students. These act as foundational currencies of AI research and are crucial, achievable artifacts for the long-term American AI ecosystem.
While many direct consumers of open models are academics, this community is far from the only group that will benefit immensely from a new wave of American open models. The low cost, flexibility, and customizability of open models makes them ideal for many use cases, including many of the ways that AI stands to advance and transform businesses large and small.
If the United States does not create its own leading open models, the focus of American researchers and businesses will continue to shift abroad. The benefits of openly sharing a technology accrue to the builder in mindshare and other subtle soft power dynamics seen throughout the history of open source software. Today, these benefits are accruing elsewhere due to the intentional support of open models by many Chinese organizations. The gap in performance and adoption will only grow as the American ecosystem sees strong open models as something that is nice to have, or an afterthought, rather than a key long-term priority.
China is adopting the playbook for open innovation of language models that the United States used to create its current AI leadership, yielding rapid innovation, international adoption, and research interest. The collapse of American dominance in AI research is driven not only by the remarkable quality of the Chinese ecosystem, but also by the commitment of China to these very same Open Model Principles - the principles that American scientists used to start this AI revolution. This is reflected further in a consistent trend of Chinese open models being released with more permissive terms of use than their American counterparts.
The many leading closed research institutions in the United States are still creating world-class models – and the work they do is extraordinary. This collapse is not their fault, but closed labs make closed research, and the acceleration of AI was built on open collaboration with world-class American models as the key tool.
As researchers, our focus is on leading the research and development for the core technology defining the future, but there is also a growing list of other urgent security and policy concerns facing our nation around the lack of strong open models. To start, adoption of open models from the PRC in the US and our allies has been slow in some sectors due to worries about backdoors or poor security in generated code. Similarly, there is concern over the outputs of these Chinese models being censored or inconsistent with everyday American values of freedom, equality, and independence. There are even parallels between how the PRC’s national AI champions are increasingly racing to release cheap and open AI models and the PRC’s historical practice of dumping state-subsidized, below-cost exports from China to undermine American competitors. With the dynamic and rapid evolution of this technology, we need to get ahead of these issues before stronger habits, cost disadvantages, or other incentives reduce the practicality of adopting American open models.
America's lost lead in open model performance
On countless benchmarks, the leading American models have fallen behind counterparts from Chinese companies. In July 2024, American models in the form of Llama 3 had leading performance over any openly available Chinese models. Since then, a growing number of Chinese open model providers have surpassed and widened the performance gap with the leading American open models.
The leading American open models are Meta's Llama and Google's Gemma models. The Chinese open models from DeepSeek and Alibaba's Qwen have traded off positions at the frontier of capabilities ahead of their American counterparts. However, the Chinese ecosystem is expanding rapidly, with new players such as Moonshot AI (Kimi), Zhipu AI, or Tencent close behind.
We consider two popular public, aggregate benchmarks to demonstrate the state of China’s current open model dominance. These represent crowdsourced rankings, LMArena, and comprehensive intelligence rankings by blending a variety of capability benchmarks, from ArtificialAnalysis. The pace of progress on these Pareto frontiers is only part of the equation. In addition to leading, the top 10 open models on LMArena are all created by Chinese organizations. For ArtificialAnalysis rankings, the top 3 open models are of Chinese origin as of publishing on August 4th, 2025.
The isolation of Meta's Llama
Meta CEO Mark Zuckerberg has been one of the few clear advocates for the long-term imperative of America building open models. Since the release of ChatGPT, this has been manifested by Meta's Llama series of models – these had long been the definitional open models that served as the basis for research and product development in 2023 and 2024. This basis for research is established by releasing a suite of strong models across a variety of sizes. The original LLaMA family came with models of 7, 13, 32, and 65B parameters, which quickly became defaults of the research community based on convenient factors of them fitting on certain popular GPUs for finetuning or inference.
For a first instance showcasing the gap in adoption, the Qwen 1.5 family of 8 models was released shortly after the Llama 2 family of four comparably sized models in the summer of 2023. An analysis of cumulative model downloads shows the Llama 2 models being downloaded about 500% of that of early Qwen models (a difference of 10M versus 60M total downloads with half of the models), highlighting the original state of play in the open ecosystem – a large lead for American models.
Llama 3 continued this trend with a series of models across 2024. Pieces of the Llama 3 family (and its various versions in Llama 3.1 and 3.2) are some of the most popular models ever in HuggingFace’s history as the leading distributor of open models. At the same time, the newer Qwen models from Alibaba, this time the Qwen 2.5 suite of 2024, showed substantially closer adoption numbers to Meta’s Llamas – a lead of only 20 million cumulative downloads for Llama 3 over the Qwen 2.5 suite with both of them crossing over 120M total downloads.
Llama’s lead was built on a combination of strong performance and existing distribution channels. This success came in spite of a restrictive license – the contract between the open artifact’s creator and the downstream user – that can require nuanced legal consideration about if a particular use-case is compliant. Meanwhile, Qwen and other Chinese models have adopted simpler licenses drawing on historical practices in open-source software (OSS), removing another barrier to uptake on their models.
Meta has effectively been a singular horse in this race. As language models were established as a core technology, competition has arrived. Between the last releases of Llama 3 and the arrival of Llama 4, the landscape of open models changed substantially with the arrival of DeepSeek’s permissively licensed, frontier models in DeepSeek V3 and DeepSeek R1. Now, Meta was effectively alone in releasing its best models regularly and expected to compete with Qwen making large families of models great at any size scale and DeepSeek releasing open frontier models. Both types of models are crucial to the health of the ecosystem, but they can take slightly different foci to get right.
China today has 5 amazing open labs, a number which is growing, and America has Meta as its open models champion. We are running Meta in a race against 5 other Chinese runners, and then complain when it doesn't win every race every time. Our problem is not Llama 4 being not state-of-the-art; our problem is running a solo athlete against a team built with an ecosystem to support its growth.
Chinese open models are taking the all-time lead in adoption
The available data showcasing adoption of open language models – how much models are downloaded and how much base models are modified for new uses – shows that China has taken the lead in recent adoption and will soon take the lead in all-time adoption.
We collected historical, daily download data from 6 of the leading open model providers across the world – Meta, Google, Mistral AI, Microsoft, Alibaba Qwen, and DeepSeek AI. Grouping by locality we can see America’s early lead with Llama, Europe’s surge with Mistral’s early viral releases almost surpassing the U.S. in April of 2024, and a consistent acceleration from the Chinese providers until they’re surpassing the U.S. this summer. As of August 2025, the leading U.S. and Chinese models both have around 300M total downloads on HuggingFace with the Chinese rate of growth being notably higher. The growth rate for European models has remained lower, with their cumulative downloads reaching around 100M today.
An important benefit of open models is the ability to finetune them, a process to adapt a given model to a specific purpose. This process is at the heart of academic research and important for businesses to shape a given model to their individual needs. While there are more cumulative derivatives of American models at the moment, Chinese models are gaining momentum, especially this year.
Early in 2024, Chinese models accounted for 10-30% of the new finetuned models appearing on HuggingFace. Today, derivatives of Alibaba’s Qwen models account for more than 40% of the language models appearing on HuggingFace month over month (the overall picture is quite similar to the downloads data) – and that is just one of China’s leading open model laboratories. Meta’s share of derivatives with the Llama models has dropped from a peak of nearly 50% in the fall of 2024 down to only 15% today. With far fewer open model options appearing from the U.S. or Europe, the proportion of Chinese models in the AI ecosystem is expected to continue to rise.
What the ecosystem needs
We can fix this. America has the talent, compute, and capital to lead open model development – we just need to get them to the right place.
The tone for change is well represented by the White House's recent AI Action Plan, which paints a much clearer vision for the benefits of innovation and adoption globally to far outweigh the current measured risks. This represents an inflection point in the perception of open models, especially in the United States, but we still have a long way to go to support this vision with artifacts and actions.
The United States has a thriving AI research community, but it is missing the models that it itself has created and has complete knowledge of in order to create clear, and rapid progress. For example, the area of research with the most excitement following recent reasoning models is reinforcement learning with verifiable rewards (RLVR). This research has largely been performed on Alibaba's Qwen models from China due to their strong performance across math, code, and STEM benchmarks.
There are two categories of truly open models that we need in order to lead on all metrics of open models defined by how AI is studied and used. Both are essential and complement each other and the rest of a leading AI ecosystem. The best outcome is when these are accompanied by training data, intermediate checkpoints, base models, training code, and permissive licenses accepted as standards for free use by the AI community. These models with everything released, currently less common across the industry, are known as “open source models” to clearly note the benefits that come with more knowledge of how it was built.
First, we need leading open models at the frontier of performance. These should be the best models in the world and can be complementary to offerings from the leading closed AI models built in America, offering cheaper costs and more modifiability. The fundamental insight driving the recent rapid buildout of AI training infrastructure is the idea of scaling laws – this applies to open and closed models alike. The ballpark of scale needed to reach the leading edge of performance today is 200 to 600+ billion parameters with a mixture of experts (MoE) architecture – a size range used for all the leading open models from the U.S. and China in 2025 that challenge the best closed models on intelligence benchmarks.
With these leading models, we need a family of related models across a variety of sizes to allow every application and direction of study to be addressed. This is a standard adapted by leading open model suites from the U.S. and China alike. Only the most challenging tasks need the largest models, and for the rest of the tasks facing AI there needs to be tools to understand the minimum model size to solve certain simple tasks. A distribution of model sizes from those that can run on your iPhone to those that are assisting with the hardest intellectual work and everything in between creates maximum opportunity to advance and integrate AI broadly.
The entry point to train models of this size distribution is a cluster of compute on the order of 10,000+ leading GPUs. It is standard for top models to be trained with small teams of fifty to a few hundred people. A famous number on the cost of training frontier AI models from earlier this year was the often quoted $5 million figure for DeepSeek V3 – this is misleading on what it actually takes to develop these models, and the authors of the DeepSeek technical report acknowledged so much. 10,000 GPUs provide an entry point for rapid iteration concurrent to large-scale training.
America should target having multiple centers producing excellent open models. This serves to de-risk progress on training these models, given the urgency of the mission, but will also allow for a more diverse set of artifacts and for the research groups to learn from each other without first making the training organizations so large that progress is slowed.
There are many avenues to obtain and allocate these resources across multiple stakeholders. We need to engage across private companies, philanthropic institutions, and government agencies. Programs such as the National AI Research Resource (NAIRR) are important for broadening access to resources related to AI research including compute, data, software, and models, but these ecosystem-wide solutions are not enough to create breakthrough models as China is with concentrated bets. We need immediate, targeted interventions that can deliver frontier open models within 6-12 months, not years.
As many organizations around the world create strong AI models, it is becoming clearer that with the right compute and talent, strong models can follow. The formula we must follow is delivering these resources with the directive to release the models openly, then we can solidify American AI leadership. Every stakeholder – from tech giants to philanthropies to federal agencies to researchers and engineers – must ask themselves: Are we funding or participating in the future of AI research, or are we ceding it to competitors who understand that open models are the foundation of AI supremacy?
I’m excited to welcome Ross Taylor back on the podcast (and sorry for the lack of episodes in general – I have a lot going on!). The first time Ross came on we focused on reasoning – before inference-time scaling and that sort of RL was popular, agents, Galactica, and more from his Llama days. Since then, and especially after DeepSeek R1, Ross and I have talked asynchronously about the happenings of AI, so it’s exciting to do it face to face.
In this episode we cover some of everything:
* Recent AI news (Chinese models and OpenAI’s coming releases)
* “Do and don’t” of LLM training organizations
* Reasoning research and academic blind spots
* Research people aren’t paying enough attention to
* Non language modeling news & other topics
Listen on Apple Podcasts, Spotify, YouTube, and where ever you get your podcasts. For other Interconnects interviews, go here.
Show outline as a mix of questions and edited assertions that Ross sent me as potential topics.
00:00 Recent AI news
Related reading is on Kimi’s K2 model, thoughts on OpenAI’s forthcoming open release.
* What did you think of Z.ai’s GLM 4.5 model (including MIT licensed base model) with very strong scores? And Kimi?
* What will OpenAI’s open model actually be?
* What do you make of the state of the ecosystem?
12:10 “Do and don’t” of LLM training organizations
Related reading is on managing training organizations or the Llama 4 release.
This is one of my favorite topics – I think a lot of great stuff will be written on it in the future. For now, Ross asserts…
* Most major LLM efforts are not talent-bound, but politics-bound. Recent failures like Llama 4 are org failures not talent failures.
* Most labs are chaotic, changing direction every week. Very different picture from the narrative presented online.
* Most labs represent investment banks or accountancy firms in that they hire smart young people as “soldiers” and deliberately burn them out with extremely long hours.
36:40 Reasoning research and academic blind spots
Related reading is two papers point questions at the Qwen base models for RL (or a summary blog post I wrote).
I start with: What do you think of o3, and search as something to train with RL?
And Ross asserts…
* Most open reasoning research since R1 has been unhelpful - because not enough compute to see what matters (underlying model and iterations).
* Best stuff has been simple tweaks to GRPO like overlong filtering and removing KL divergence.
* Far too much focus on MATH and code - AIME has tens of samples too so is very noisy.
* People are generally building the wrong kind of environments - like puzzles, games etc - instead of thinking about what kind of new capabilities they’d like to incentivise emerging.
50:20 Research people aren’t paying enough attention to
The research area I hear the most about right now is “rubrics” – a per-prompt specialized LLM-as-a-judge to replace reward models. SemiAnalysis reported OpenAI scaling this approach and lots of great research is coming out around it.
I start with: What do you think of the state of RL scaling and generalization? What of models losing
Ross asserts…
* Rubrics are underhyped on social media - they were driving force behind projects like DeepResearch - and GenRMs are interesting but perhaps slightly overhyped.
* There is an evals crisis - there are not enough high quality evals, particularly for frontier tasks like automating research and real life work. Impediment to anyone building agents or ASI.
01:02:46 Extra stuff!
I ask Ross: What AI are you using today? Why?
To conclude, Ross wanted to discuss how AlphaEvolve has been underhyped on social media, and means the future isn’t just RL. Shows there are other effective ways to use inference compute.
Interconnects is a reader-supported publication. Consider becoming a subscriber.
Transcript
Created with AI, pardon the minor typos, not quite enough time this week but I’m hiring someone to help with this soon!Nathan Lambert: Hey, Ross. How's it going? Welcome back to Interconnects. I took a many month break off podcasting. I've been too busy to do all this stuff myself.
Ross Taylor: Yeah, I was trying to think of all the things that happened since the last time we did a podcast a year ago. In AI time, that's like two hundred years.
Nathan Lambert: Yeah, so I was looking at it. We talked about reasoning and o1 hadn’t happened yet.
For a brief intro, Ross was a co-founder of Papers with Code, and that brought him to Meta. And then at Meta, he was a lead on Galactica, which was a kind of language model ahead of its time relative to ChatGPT. So if people don't know about Galactica, there's a great paper worth reading. And then he was doing a bunch of stuff on reasoning with Llama related to a lot of the techniques that we'll talk about in this episode.
And now he's doing a startup. I don't know if he wants to talk about this, but generally, we talk a lot about various things. This got started through o1 and trying to figure out scaling RL. We started talking a lot but then we also just resonate on a lot of topics on training language models and other fun stuff - and also trying to be one of the few people not in these big labs that tries to talk about this and think about what the heck's going on. So we're gonna kind of roll through a long list of a lot of things that Ross sent me that he wanted to talk about, but this will be a compilation of the things that we've talked about and fleshing them out outside of the Signal chat.
So, Ross, if you want to introduce yourself more, you can, or we'll just start talking about news because I think a lot of people already know you.
Ross Taylor: Yeah, let's get into the news. There’s lots of fun things to talk about.
Nathan Lambert: So, the last two weeks of Chinese models. I think we had Z.ai's GLM 4.5 today. Kimi-K2 last week. I think Qwen is on a roll. I thought summer was supposed to be chill but this is crazy.
I haven't even used all of these. The pace is just incredible. And all the open models have actually good licenses now. But is this going to hurt anyone in the US? Where do you see this going in six months?
Ross Taylor: Yeah, so yesterday was the one day I actually tried to turn off Twitter. And so when you told me in the morning about the new GLM model, I had to read up on that. So that shows if you take your eye off Twitter for one second, then you’re behind on open source...
But yes, I think the general theme is that it’s been absolutely relentless. So thinking about the last time I spoke to you on the podcast a year ago, Llama 3 was a fairly established standard.
There were still things happening in the background, if you paid attention to things, but now it's absolutely relentless. In the case of China, I think their business culture is that - as soon as they find something is successful - they’re very good at concentrating resources and going after it. So it’s created a very competitive space.
I think the context is very interesting in several different dimensions. There's the geopolitical dimension, which you've hinted at in some of your blogs. For example, what does it mean if the open source standard is Chinese? What does that mean if we think about these models not just as things which power products, but as (critical) infrastructure? Then it seems like China has a great advantage if they want to be the standard for the whole Global South.
Nathan Lambert: Yeah. There are a few things that we're going to come back to in this conversation that are so interesting. We're gonna roll into what it takes to train these models. And we're going to talk about how crazy, political and hard it is in the US. But we have all these orgs popping up in China - so is this partially just a US problem?
But then we also have OpenAI that's supposedly going to release a model. There are multiple things. But my question is: why is China doing so well? Are they well suited to training these language models?
Ross Taylor: I’ll caveat what I’m about to say by saying that I want to be careful about making generalisations. Because, for example, we’ve seen some of these new Chinese organisations be good at innovation - for example, this week we had GSPO which was nice. But for Chinese orgs, my general sense is that, once something has already been validated, the specification for what to build has been set, and the task can be reduced to an engineering problem, then Chinese culture is very well set up to succeed in those situations.
The other dimension which has become relevant - especially after DeepSeek - is that the Chinese Government has traditionally been very good at recognising what’s successful, pouring resources in, and facilitating public-private collaborations. I think that surprises people still in the West. For example, people are surprised that a group can come out of Tsinghua can and fairly quickly have their own state-of-the-art LLM. Why isn’t there a similar story for groups coming out of MIT?
Nathan Lambert: I’m not sure about this.
Ross Taylor: I think the US will eventually wake up to this, but…
Nathan Lambert: My understanding is that Z.ai is a startup that spun out of Tsinghua, so I don’t know if it’s the best comparison. Also Alibaba is the clear winner here because they have Qwen, but they’ve also invested in Moonshot, which is Kimi, and then I think also Z.ai.
So I’m more interested in the question as to why they are all open. That seems more important relative to talent because there are lots of universities that might have model orgs spinning out of them - even in the US - and it’s not solely a Chinese thing.
I think it could happen with a group out of MIT. That being said, I agree that the US should have more compute deployed for academics and a lot of universities are just spinning them up now. It just takes a long time.
So I think there’s a lot of things that Twitter is mixing up here. There's a good tweet in it, but I don't think it'll be 100% true, which makes for a very viral tweet when it feels true.
Ross Taylor: Yeah, I think there is definitely naivety about how things are actually working (in China). And there’s asymmetric information, in that you don’t truly know what’s going on in the inside of these organisations.
The other thing worth mentioning - which is maybe a separate topic - is that there’s a tendency to see open models as a homogenous category. But there are very different use cases. So if I want to do a new reasoning paper, I’m going to use a Qwen model. But then if I’m doing distillation, I’m going to use DeepSeek or Kimi.
This discussion also relates to OpenAI’s rumored open model: because in my mind I still don’t quite see how it will fit into the ecosystem. Because is it going to be something that people build research on? If it’s a post-trained model, then probably not, right?
Nathan Lambert: Yeah. But their tweet was about safety, so I doubt it is a base model if they’re delaying it for safety. I do think they actually delayed it for this reason. It’s very much in OpenAI’s culture. But I don’t think it’s going to change the ecosystem. It will be an interesting one off.
I also don't expect them to release a model that's based on their GPT architecture. My bet is they take an off-the-shelf architecture like Qwen or Llama. A lot of the recent OLMo models are very Qwen-y. And they will also be deciding sizes based on what fits on what cluster - e.g. Qwen is very deep rather than wide, and OLMo 2 is very similar to that. So I think the OpenAI model is going to fit that mold.
Ross Taylor: I think so. I guess one way to think about it is they're just trying to “distill” their RL infrastructure into weight space, right? As opposed to publicising their (internal) architectural choices.
But back to the discussion, and maybe this is a question for you Nathan, but do you think their model is going to be more comparable in use case to a Kimi or DeepSeek? Or is it more similar to Qwen? Or is it actually something completely different, like an on-device model? A smaller model?
Nathan Lambert: I expect it to be smaller. They joked about on-device, which I don't know is the right framing.
Ross Taylor: Yeah.
Nathan Lambert: I'm also just now realizing how - if RL is their great strength - then part of the challenge of shipping an RL model in open source is that you need your training infrastructure to match the inference infrastructure. So unless they train this on an exact VLM that people have access to - and some open source environments - then they can’t just dump the model and expect people to be able to do search and code execution in the open model stack.
I don't know exactly how Qwen and DeepSeek have gone about this. My impression is that they're actually not as useful in terms of tool use because it's so hard. I think that tool use is naturally a closed model reinforcing thing because it benefits to have these tools match up.
Ross Taylor: So the Qwen models are pretty good at things like function calling. Kimi - at least in the benchmarks - was also pretty good at agentic tool use benchmarks. And then - this is a separate discussion - but they had this nice training innovation where they use lots of MCP servers in a synthetic data strategy. But then again, you’re mostly seeing indications of capability in headline evals, which you shouldn’t really trust anyway.
Nathan Lambert: I think of Claude 4 as the release that ended eval chasing. On paper the release was so lame, but it delivered for everybody - which is very bold because there is a lot of money on the line. They are constantly fundraising and if one fundraiser gets spooked because the release numbers are bad, then that’s a lot of CEO calls that they have got to make.
Ross Taylor: On evals, I was thinking about this a few months ago. It might have changed now given the pace of AI development, but I was thinking about how you might split up the impact timeline for a release.
So day one is headline benchmark numbers - which are mostly b******t. Like I’ve got this amount for my model on MMLU Pro. But then the next tier of impact is the day after the release where people have all these weird bespoke evals on Twitter.
Nathan Lambert: The pelicans and the rotating hexagons and balls…
Ross Taylor: Yes, and by this stage you’re getting more confident. Because unless the model developers are very smart (which some of them are), then they probably haven’t optimised for day two benchmarks. So at that stage you’re beginning to believe that the model actually generalises beyond the headline numbers.
And then finally you have a week or two weeks after the release where you can say that you’ve tried the model quite a lot now, and you then have real confidence that the model is good.
Nathan Lambert: Yeah. Refute my claim: Chinese providers are still optimizing for benchmarks more than OpenAI, Google, and
Ross Taylor: Yep, I mean it’s probably true.
Nathan Lambert: It feels so obvious to me. I think that China has closed the gap to a remarkable degree, but I don't think they've caught up fully. I think that's hard. It’s very hard to get all the data and pipelines in place. A lot of it is actually user data, knowing your user, and hill climbing that. So for example, all these APIs not working is a huge issue for them.
Ross Taylor: Yeah. I think (Chinese models) have also been helped by the fact that a lot of the academic work that builds on them has been doing reasoning work in publicly available data domains like math and code.
The models have been heavily optimised for these domains anyway, so the model developers are not quite as exposed - since people aren’t really testing the true generalisation capabilities of the model. We already know that the Qwen models are heavily mid-trained on math and code, so they will hold up performance-wise there.
Nathan Lambert: Yeah. Okay, this is a good preview for the episode. I think that the main things are going to be how to build good organisations, and then academic reasoning research and how to bridge the gap. I think we can talk starting about org charts.
So how do you make a good org? Or maybe there are two things. One: how do you make a good org chart for training language models? And two, how do you make an effective culture?
I think this is quickly becoming one of my favorite little niche interests because there's just so much intrigue in it. There's just so much money on the line to break everything. So you sent me some hot takes if you want to read them, but the floor is yours for what doesn't work.
Ross Taylor: Sure. So if anyone’s been on social media recently, the general trend nowadays is to check your phone and see these NFL draft style tweets about researchers moving between orgs.
First of all, researchers have always moved between orgs. This is not a new thing. And a lot of the org moves that were talked about - at least outside of Meta - were just regular moves.
But I think the bigger mistake on Twitter is just the tendency to see the bottleneck in LLM projects as skill issues. And at least from my n=1 experience, that has never been the main bottleneck for success.
There are a number of ways to make this case, but I think I'd start by saying that machine learning is a heavily empirical science. So what does genius mean in that context? What does talent actually mean?
There are certainly some skills which are useful - like how do you form the right minimal viable experiment? And how do you iterate fast to explore a research direction where you’re going to hit a lot of dead ends. But a lot of it comes down hard work, good infrastructure, and ultimately resources.
So in that context, most of these orgs - even before public failings - had very good people. And I don’t think the difference in talent between orgs is that large. Smart people will eventually figure things out. So therefore, more often than not, the difference between a good versus a bad model is reflecting an inefficiency in the ability to channel resources to your talent. And that is the fundamental point.
Now you could say, on the flip side, okay, Ross, well, if that's true, why is Zuck paying people these massive amounts of money? And I think that's a separate question. But yeah, more often…
Nathan Lambert: Well what do you think?
Ross Taylor: I am torn on this because, on the one hand, I think the new group will probably make very good models. They’re very smart people. And I think having a new org as well is the right way to do it.
I think in leadership's mind, it's a case of “Look, we tried this multiple times, we’re very serious about this, we have resources, so let’s do the maximum conviction play”. And I think that's broadly what you should do because it’s a big expense, but it’s not massive, massive spend (for these large companies).
But on the other hand, I feel sorry for - this isn’t a Meta point by the way, but a general point - but I feel it’s a shame these organisations don’t have good mechanisms to identify the talent they already have in their orgs and have to recruit externally.
The talent that has already done the hard work, that is. It’s a shame they have to hire externally and start afresh. That’s the tragedy.
So that’s the conflict in my mind. I think they’ll make great models. I think it’s the right approach to do things afresh. But at the same time, it’s a shame that all the people that came before them, and made the previous generation of models, are treated like an asset. In the sense that you’ve used these people - grinded them really hard - and now you’ve moved on to a new group of people.
Nathan Lambert: You put this in your provocations. You said LLM labs are like investment banks where people are slotted in to burn out and burn through. I know that a lot of the work that needs to be done is somewhat mundane data work and it can be parallelised - e.g. if your users are asking this type of question, let’s create new prompts and manage human works and create synthetic data pipelines. And that works a lot of the time.
But then, I remember the Dwarkesh podcast with Sholto and Trenton - and it’s the one where they’ve both moved jobs (which reinforces your point), but they were saying you just need to convince someone at a frontier lab that a particular problem is important. I.e. people talk about things, but they just have to do it.
So is it the case that people are just dispatched to solve specific problems, or do individuals have free rein, and it’s fun on the ground because you choose the things you want to add to your beautiful final model?
So you can present a positive and a negative. It might vary across labs, but I guess your provocation is that there's a bunch of places where it is a meat grinder and you just put people in and chew through them.
Ross Taylor: I think so. Unfortunately the model for a lot of successful tech companies is to get very young, motivated, people - with a base level of intelligence - and make them work very long hours on a project with a big mission. This was the classic Elon way to run a company.
But this is also the model for a lot of frontier labs. You have your soldiers who - on the surface - look similar to quants at hedge funds from like 10 years ago in terms of their working hours. And in the culture too, you have friendly competition between people who all want to be the best.
Nathan Lambert: I will say: I know a bunch of people at OpenAI, and they do work crazy hours. I also work a lot, but I do a lot of things that aren't grinding data to go into the model.
Ross Taylor: Yeah, so on the question of decision-making, I think major decisions are generally made by people who are a little more experienced and already have some successes to their name. But you do need to have soldiers in this kind of environment. The space is just highly competitive (and requires people to work long hours).
And I think that's a shame. Even for myself right now, where I’m trying to build a startup, I’m thinking that - yes, we all need to work hard - but is there an alternative model where you invest in your employees instead of using them? - i.e. burning them out and then moving on to a new group. That’s what I’m trying to work out for my new company.
Nathan Lambert: I feel like a lot of people are just more cynical now in tech, myself included. I got a great cold e-mail from someone fresh out of undergrad, and I was pretty sure in two to three years this person would be legit. And I was talking to a coworker on how we could potentially capture this and invest in them. And they were just saying we might get them, but then they’d just go to OpenAI in 2 years. So we don’t get any of the upside.
I think some of that is just cynicism. Investing in people is still the right thing to do because you’ll end up keeping the ones that are a bit more grounded even if it is really hard. For example, I've lost people that are extremely talented that I wouldn't want to keep. So I don't know how to balance that cynicism versus reality of building teams in the long term.
I guess smaller teams might be a bit easier to maintain, whereas if you’re at a tech company, the churn is hard to avoid because there’s so many levels in moving up.
I think some of the rumors around Meta and Llama 4 - at least from the Dylan Patel SemiAnalysis article - were about them doing these cowboy crazy model training runs, including changing pre-training mixes half way through, and that maybe points to dynamics with middle management wanting their data to be used so they can get promotions. But most labs I don't think are doing that type of s**t for their leading models. And I don't think Meta is normally doing that. I think that was a pressure cooker side effect.
Ross Taylor: I would push back on that a bit by stating that all of these labs are deeply chaotic places (not just particular orgs). They change direction every week, right? That’s just the nature of the field we’re in.
But then, it is definitely true that certain labs are good at projecting, at least externally, that they have their s**t together. They have AGI internally, all this kind of b******t.
The truth is that it is a shitshow everywhere. It's just that if you're going to be a s**t show, you at least want to be a functional s**t show, and you want to make good models. Right?
As I mentioned before, I think there are new plays to be made around taking the view that you want to invest in your talent as opposed to just grinding them out. But I would also say that, in lab culture, people tend to overvalue raw talent again - especially in empirical science. If you take the view that an empirical science is mostly about experimental velocity, then you don’t just value infrastructure in that world, but you also want to hire folks who are very collaborative and who want to help each other.
It sounds like a b******t point in a field that lionises individual intelligence, but I just feel that if you're making a marginal hiring choice, then you have to think about how someone adds to an existing group? So I think there are new plays to be made on talent.
But there is nuance. Because there are certainly people who are especially productive. I’ve seen that in person. So it’s not like everyone is equal - that is definitely not the case - but I just feel that individual talent is overemphasised when problems in these orgs are mostly structural.
Nathan Lambert: The differentiation right now is people who are willing to put more highly focused hours turning the crank. Every organisation has the baseline time costs of needing to do meetings, commute time to work, commitments etc. But in terms of AI, where people are doing more and more, this really favors young people who don’t have a lot of responsibilities.
Ross Taylor: This is maybe a transition onto another topic, but I’d make a more controversial point which is that - even the things in ML which seem more like novel research are more the result of persistence rather than inspiration.
For example, this time last year we were both speculating about what o1/Strawberry was. And speculation makes you think it was some amazing new thing. But actually it was basically what we were both doing two years ago right? Essentially RL from verifiable reward, but with very good base models, because they were in a good position to exploit that, and then enough ablations to find a recipe that worked.
So this is oversimplifying things a little, but we should take the view that they just had to do the work to make the recipe good. And that comes down to experimental velocity, and also having the right infrastructure and a good enough base model. So in that world, what is talent?
Is talent the person who says “we should make the models think more”, or is talent the person who is actually on the ground doing the ablations to find out which recipe works? Right? Because I can also make models think more by best-of-N, but, then there may be better ways to do it?
Nathan Lambert: I mean, I think I analogize a lot myself with my athletics career - like rowing in college. I think so much of it is the same. I wasn't the most gifted athlete, but if you put in the hours and you understand where you're spending your effort, it works out for people.
The question I wanted to ask you on this topic is, given that that these orgs are so chaotic, then what does this mean for the ceiling on progress? One of the most coveted questions is about the trend line. There are obviously going to be new paradigms - inference time scaling was an obvious one if you thought from first principles about what compute and intelligence is - but even if we don’t have a new paradigm, then what is the ceiling?
Ross Taylor: I would say that, even in climates where most organisations are chaotic, you’re still going to have macro factors that lift all boats. So a good example recently was these gold medal results on IMO. Three or so different labs all had different approaches and all found they crossed the threshold for a gold medal.
If you were to zoom out - and one way to do this is to imagine you're looking twenty years into the future back at this time - then would you look at the individual methods that researchers used, or would you just say compute reached a critical threshold where things began to work?
So compute is the big exponential that's underlying all of this. And then if you zoom into a shorter time horizon, then you're seeing more of the local challenges, like what’s the particular bottleneck at a point in time? So maybe the bottleneck to agentic models is scaling RL environments. Or maybe the bottleneck to better reasoning is longer context windows.
But look: fundamentally as long as compute keeps coming online, I think the trends look good and all of the organisational chaos is short-term noise. It slows down progress a bit but is not meaningful in the long-term. But, unfortunately, it's still meaningful for people in their careers because one to two years of organizational chaos could matter personally. But on longer timelines, it doesn't really matter.
Nathan Lambert: Yeah, I agree. It seems like the question is what happens when the fundraising starts to slow down. We're on a trend line of compute rollout. But if Sam Altman can't raise again, that is a very big sign. That's like the end of the “bubble”. OpenAI is not going to go away because of that, but if they can’t get the next cluster… then that would be a bad sign.
Ross Taylor: I'm quite optimistic because I think you only have a bust if AI ceases to be increasingly useful or doesn't live up to certain promises. But even if there's no algorithmic progress, I still think AI will continue to continue to be increasingly useful. I don't think there are fundamental barriers. It's just a question of how quickly you get things right.
I think the argument would have been slightly different two years ago. If the reasoning paradigm didn't come through, then I think it would have been trickier to justify the expense because then you'd be looking at reasoning benchmarks and thinking: s**t, to push this forward I need this amount of data annotation or need to generate this amount of data.
Nathan Lambert: You look at GPT 4.5 as the example.
Ross Taylor: Yeah, exactly. That's a really good example. So you can treat that model like a counterfactual universe where reasoning didn't happen. There we would all be looking at the model thinking “it's good at creative writing, but maybe not so good at some more things we really care about (like reasoning)”.
By the way, I'm sure it’s a really good model. I didn’t play with it enough to form a good judgement.
Nathan Lambert: I've been using it a lot. I used it for a long time - especially until Claude 4 - as it’s just nicer, especially when GPT 4.1 was so sycophantic. But GPT 4.5 was nice.
Ross Taylor: So I'm gonna flip things around and ask you a question Nathan. Let's say we are here in a year's time. What does the key benchmark look like for LLMs that everyone is focused on?
Nathan Lambert: Oh, it's fully gonna be some agentic thing. I don't know if it'll be as stupid as making money on the stock market… I wrote a post on what I thought was coming next. One of the most poignant things I was looking at is the fact that scaling models is no longer the direction anymore. All the marketing is shifting to agents. And I think some of that is because it's not easy to scale parameters anymore.
Every RL curve is this log plot, and it becomes hard. But agents are already beginning to work well. For example, this year Claude Code showed up. There's gonna be versions of that in all sorts of domains and more people working to evaluate them. That will create an interesting marketing problem where labs need to figure out how to communicate that their model is good.
But the future looks like it’s all on the agentic side, and will lead to a big shift in what the language modelling companies need to think about. The prioritization of the company is also different, whereas modelling was always central before. I’m still modelling-pilled and think that is the central thing for the company…
But it’s true that now that teams building products are going to hold more weight than they used to. And there will be interesting changes in how these companies manage this transition, and how communications change.
So, I think Claude Code is great. But I think that it's hard to integrate in some things. For example, how do I get that running on my cluster at AI2 where we have all of our data and models, launch evals from our file system on the GPU machines. I don’t think that quite works yet, but maybe I’m doing something wrong.
Ross Taylor: Yeah, I agree with your answer. So I spent several years working on Papers with Code, where we were trying to focus heavily on evals before they were a big thing - trying to index all these various leaderboards. And I think now is an interesting situation because I feel like if you make good evals now, you possibly have more leverage than you've ever had in the field of ML..
This is a weird thing because traditionally evals were quite an unsexy thing to do. It was a thing that researchers didn't want to do because they'd rather be training models. But now the ability to define a metric for a capability that you'd like to see - e.g. trading stocks, or doing scientific research - is just incredible leverage that you can wield. A small group of people in places like universities can say “this is the new north star that we should achieve for agents” and shape how AI progress evolves.
Nathan Lambert: It can happen. We recently released IFBench, a benchmark for following instructions which is just more constraints and a different prompt sourcing. And I was saying to folks that we need to have the goal of making at least two frontier labs adopt it. And I messaged various people, including someone at OpenAI, and they said they already integrated it last week.
So yes, someone doing research (on evals) has a shot at getting into the OpenAI internal evaluation platform.
Ross Taylor: Exactly, so it's incredible leverage. And then the other interesting thing is that the friction for making and using good evals is going to increase quite a lot.
For example, in some of the recent benchmarks, you need the RL agent to have access to a GPU and then you need to spin up lots of these servers to do rollouts. This is expensive. Long gone are the old days where you had two CSVs with a train and a test split.
And then on the eval creator side, there’s a big difference between good and bad evals as models become more capable.
A bad eval just means that you're going to get incredibly egregious reward hacking, and you're not going to learn anything useful, whereas a good eval is a pathway towards a brand new capability.
Nathan Lambert: I have a related question on this. So I see three eras in evals based on what people are doing with models.
For pre-training, the best evals are testing knowledge and these very broad things and are hard to game. It's just kind of like FLOPs.
At post training, a lot of evals are formatting and extraction. I think formatting became even clearer to people when these RL environments became the hot new thing. And I actually think that post training might be like the ugly duckling in the middle, where then if you go into agents, all the agentic tasks are gonna be evals of actually doing things and you can't like format-lie your way through that. So it might be that post training evals are the hardest one to get right.
Ross Taylor: Yeah, and I think you're going to see more cases of people claiming good results, but when you look beneath the surface, you’ll see insane reward hacking. So the meme right is KernelBench evals. Have you seen these?
Nathan Lambert: Oh.
Ross Taylor: You see all these amazing speed ups which aren’t even possible based on the hardware. And this is not a problem with KernelBench, I would say it’s more a problem with people publishing papers for agentic evals and not looking at their results carefully.
So this shows that to get an eval in the right place takes a lot of work. And even with progress in models, I don’t think you’re going to be able to fully automate the construction of a good eval in the next year at least. I might be wrong. Models will certainly help us in creating evals. So I think that, for now, it’s a place where a researcher can have a lot of leverage.
I think if you were to ask what is the central eval is right now, it'd probably be something like SWE-Bench (verified). But even that is now quite saturated. So there's a big blue sky now where someone can define what the next big task is for ML. And you don’t need a big cluster in order to be the one who defines it; so I think that’s quite exciting.
Nathan Lambert: Yeah. And when you think about the amount of money that'll be steered by these things, it's so crazy to have the uncertainty there and like who will come up with that as well. I think that it's part of what makes it fun, I think.
We should talk about reasoning things.
Ross Taylor: Reasoning. Yeah.
Nathan Lambert: Where do we start? I don't think I've ever done that much of a rant about the academic community chasing these things. I understand why academics are claiming to do new algorithms that get remarkable scores, but a lot of these papers are just extracting things that are hard to document from a model or something else or formatting
I was on one of these papers, which was hilarious. We figured out that if you train Qwen on random rewards, the evaluation scores go up. And we had to go through the logic on why this can happen.
Because if there's no reward, the advantage is zero and the gradients are all literally zero. And then it turns out that the algorithm manipulates the most common sequences. It's actually something that if you read a lot of the reasoning literature, people talk about how we want to make sure our algorithm doesn't squash uncommon sequences. And then the real hammer is that, if you do random rewards, then you see that the model has modal collapse onto the things that it was trained on. And that can make scores go up.
So if you have a model that two thirds of the time has a certain behavior in its reasoning and that behavior is good on the benchmark, then just by fiddling the weights a bit then it does that behaviour more. This points to a structural failure.
I would also say it is a good example for why people should be using truly open models for research purposes and why they're so good for innovation. For example, if we knew what goes in Qwen data and if someone just filtered it and it was like, oh, look, I found the found the GPQA prompts in it…then we know data contamination has happened.
The Qwen case is borderline - I don't know how exactly to characterize it because the Qwen models are fantastic - but there's so much research that is showing that they are very likely to be doing some dubious things in terms of benchmarks. It's hard for people that aren't super in the weeds to hold both of these possibilities in their brains.
So I don't know. What do you think of the last six months? Have we actually made any progress? Has the academic community made any progress?
Ross Taylor: I think there's been little progress. I mean that in the literal sense: there’s been some progress, but it has been little. I think I can answer this question in several ways.
So after DeepSeek-R1 came out, there were two approaches in open source more generally, which was either you go down the distillation route or the RL route to make interesting small models.
The initial thing that was undervalued - at least from an engineering perspective - was that for smaller model sizes, it is far more efficient to do distillation than RL.
Nathan Lambert: And not just in compute but also in performance? It's hard to do RL on the small models.
Ross Taylor: I think this point has been made twice now. So there was the original DeepSeek-R1 paper, and then more recently, there was a new Qwen paper as well. The Qwen paper showed that RL needed 17x times more compute than distillation.
So one way to think about this is that RL is a brute force lever to do data generation. But assuming that RL is still good, and you still want to do research on it in academia, then you run into a classic problem. And that problem is: if you don’t have enough compute, then you don't know if the structure you are imposing is gonna generalize (to high compute settings).
And my worry is that a lot of the results are on relatively low compute budgets, both in terms of the underlying base model, which determines how well the RL approach learns, but also the total number of RL steps. So it's just quite hard to see - unless there’s a massive gain - what’s truly important.
So the most useful things are - in my opinion - quite boring things. Like, there was the DAPO paper which showed that you should have filtering for overly long sequences, and you shouldn’t overly penalise them if your context window gets cutoff.
There has also been interesting work showing that even simpler approaches (than GRPO) might work, where you remove clipping. So Reka was doing lots of good work using REINFORCE leave-one-out (RLOO). But even there, it’s difficult because you don’t know if simpler algorithms are going to work with long agentic traces.
So it’s not clear. I think the recent work this week was actually quite good. The GSPO work was good, and if you saw their graphs…
Nathan Lambert: Explain it to people. I think a lot of people have heard of the other ones by now. But GSPO is group sequence policy optimization with Qwen Coder. Why are you positive about it relative to the other ideas? I think GSPO is well motivated but why is it getting hyped more?
Ross Taylor: So I hope I don't botch this because it's the morning. But, essentially, with GRPO, you assign a reward to the whole sequence (via the advantage). But you also have an importance weight, which is your policy likelihood relative to your old one. Because when you do RL, you typically sample lots of rollouts but do several mini batches for your gradient update. So that means you go a little bit off policy.
So to fix that you have an importance weight term. But in GRPO, while the advantage is uniform across all tokens, the importance weight is particular for each individual token. And the importance weight is calculated for a single sequence. So one way of looking at this is that, if you had more sequences to calculate the importance weight, it would be a lot less variance - but by calculating it on a single sequence, you introduce a lot of variance through that term.
So the short answer of what GSPO does is that, instead of looking at a token likelihood, they look at the likelihood of the whole sequence. So now the clipping is not on an individual token basis, but, it looks at one of the sequences in your group and says okay, this one is less likely, so we’ll clip out that sequence. And the TLDR is, at least from the results they show, it seems to be a lot more sample efficient.
I mean, it's not just 0.5 percentage points or something like that. But I think the reason I trust it more is that it’s very simple. And it’s quite directionally well motivated from just a basic understanding of importance sampling. If it were more complex, I'd be a lot more skeptical, but it's fairly simple and it seems to work well.
Nathan Lambert: Yeah, I'm still fairly skeptical.
I think academic research is relatively wide in what people are trying out but labs are relatively narrow. And once you’re further along in your modelling journey, you’re dealing with different parts of state space and then all these algorithmic tweaks just like help your model on whatever blocker it was or your implementation.
I thought for GSPO the sequence thing was funny because when you read the GRPO paper, you were like oh, the reward is just per sequence. But all the tokens in the sequence get the same loss function. But the standard implementation is to break it down per token. And then GSPO is essentially to take that standard implementation and you change the weight on every token back to this. And I was doubting whether this was really going to be a major thing.
I think for junior researchers, one of the good things about this era is that you can really learn the math by studying all these algorithms and thinking about how they are implemented. I hadn’t done that for a few years until writing this RLHF book on policy gradients and I was getting into the weeds like per-token loss, length bias for GRPO, and so on. For students to be able to do this in their brain, it is really good for thinking about the interface between algorithms and systems.
Ross Taylor: It’s interesting, because as AI became more hyped after ChatGPT, you have more people reading papers. This is a great thing, but also you have lots of new people reading papers in the wrong way.
For me the basic logic (for reading papers) is as follows: what’s the reported gain of the paper and how much complexity does it introduce?
So if you get a gain but the paper introduces shitloads of complexity, it's probably not going to stand the test of time. Whereas if it's something relatively simple, but it seems to get a good gain, then that’s the thing that is going to last.
Nathan Lambert: The o1 lesson. The simple thing. In RL research, I've heard it described as: if you see something that only beats the baseline by a few percent, it's not gonna work. But if it’s 2x then that’s a real innovation, because whether they finetune their baselines or not, they’re still going to be crushing it.
Ross Taylor: Exactly.
Nathan Lambert: So I think that's a good heuristic for people right now.
Ross Taylor: And I think researchers are their worst enemy because they want to see their own methods work. But the weird thing in ML is that neural networks “want to learn”. So if you push something enough, it will work. It's just a question of whether that is a good use of your time?
So the question is: what's the right thing to scale and push on? So that’s why - when you read papers - at least what I say to young researchers is that you should always judge how much complexity the paper introduces, and whether you trust the gain.
And then based on those three factors, you can judge whether it’s worth caring about the paper. But I can see why - if you’re new to reading papers - why you might be attracted to complicated, new techniques in papers that seem methodologically interesting.
Nathan Lambert: And researchers often manipulate the results of their peer methods in the way to tell a convincing story. And I think these algorithms are a perfect example of trying to tell a story.
Ross Taylor: Yeah.
Nathan Lambert: So when you think of cognitive behavior of paper authors, you have to take that into account too.
Ross Taylor: The other point I’d make is that - in the reasoning trace - I understand that everyone has to focus on math and code, because that’s where the data availability is. However, if a paper comes out and it’s just flexing on AIME and GPQA then that is just very uninteresting to me - and much more so than it would have been in February.
Nathan Lambert: I think code can be much better but it's hard to benchmark it. Describing what a good coding model is would take me an extremely long document.
That's not what the academic papers are doing. It would be great to have more benchmarks on that.
Ross Taylor: Yeah, and even the established ones have issues. For example SWE Bench has a very large proportion of issues from Django (so it’s not exactly representative of all software engineering). That’s not a burn towards SWE Bench - which is a great benchmark - but…
Nathan Lambert: They already won. They can take it - they won!
Ross Taylor: But, yes, it shows that there is a lot of detail to get right in making a good coding benchmark.
Anyway, it’s difficult because I am in this position where I can say - on the one hand - papers are just hill-climbing particular math and code benchmarks, and that is fundamentally uninteresting to me. But at the same time, I sympathise. Because there are not a lot of good open reasoning datasets in the open. And those that are open, I don’t think that they’re even going to be good for testing RL necessarily. They might test something more knowledge based, like medicine or something like that, which is less inference-time scaling bound.
Nathan Lambert: This could be a good time to transition. What is the status of RL scaling and generalizing? What is the status of RL outside of math and code? I think my prompt is: what do you think about o3-like models with this crazy search behavior and multi hop execution?
Ross Taylor: Yes. So first of all, I think it was greatly overstated that these models don’t generalize beyond math and code. I think what happened in practice is that, at least from what I know, OpenAI originally was very focused on math, logic and puzzles. And then eventually they had to broaden out because it was kind of too nerdy and biased towards these kinds of tasks.
But I don't think there was ever a question about their generalisation to other benchmarks. You could see that very early on. The way I think about this is: we started with math and code because it was easy to verify. And then through applying RL to those domains, models learnt certain strategies like “I shouldn’t answer early”, “I should check my work”, or “I should consider alternatives”. And at a very high level, if you just have a model that thinks for longer and checks its work more and considers more things, then that's gonna be useful for things beyond math. And that's reflected in the benchmarks.
That being said, if you want to get superintelligence outside of math and code, then yes, you probably need more specific benchmarks and datasets for that. So there the question is less about whether it generalises beyond math and code, but how far can performance go? And that’s when you get into interesting questions about, e.g, if you don’t have a numerical answer or whatever, then how do you verify things.
So rubrics are all the rage right now, but then there's also other directions like…
Nathan Lambert: Rubrics are so funny. It’s funny how they needed to be reinvented. Rubric is a funny name because it just seems like question-specific LLM as a judge. It's the most basic unit of evaluation or feedback.
Ross Taylor: So I think this was something that wasn't very covered in the open. So the reason why it became popular was that DeepResearch was the trigger. The rumor at least was - at least for OpenAI - they didn’t need many examples to do well in these kinds of research task.
It wasn’t tens of thousands of rubrics - it was probably in the 1,000-2,000 range of well-crafted rubrics for questions. But it clearly worked very well to teach a model how to browse the internet and synthesize knowledge. There's obviously infrastructural detail as well.
Nathan Lambert: What would a rubric look like for deep research in this case? For an essay it might be that the rubric says that an answer should be free of typos, have a clear argument and a good conclusion. It would have different checklists. But the DeepResearch example is more complicated and you might need to draw an example.
Ross Taylor: Yeah, so there are different themes you could have. It could be the general style of the answer. It could be - let’s say we want a review of the latest and greatest RL algorithms for reasoning - then there you might have a high level rubric saying that the answer should compare different methods, cover underlying algorithms, mention policy gradient, PPO vs REINFORCE and so on.
But then you might have, like, more detailed things where you just have a strong conviction on what a good answer looks like. For example, a review of RL for LLMs right now might include GSPO as of this week.
So rubric-based grading comes down to a list of checks, but the goal of that form of evaluation is that you’re trying to get a nice, continuous rewrad for the model to learn from - as opposed to something more binary and sharp. Because while 0/1 rewards might work okay for mathematics or unit tests, it would work less well for a task like making a good literature review on RL. The reward structure isn’t binary there.
Nathan Lambert: So how do you think of grader functions? I’ve thought about this for code, like the percentage of unit tests that pass. But then the model might just get the easy unit tests. So will reward shaping be here to stay or will it be washed away in the ever growing sea of compute?
Ross Taylor: I think it'll be washed away, but I think in the meantime, there's a lot of value in making very good handcrafted evals. And I hate the word taste, but there is still taste to begin with.
And I think a lot of these things are quite codependent, because to make a good rubric for a deep research task, then you need something that needs the ability to do deep research. If we were to say what makes a good literature review on RL right now, then that knowledge wouldn’t be in the weights of a language model - the model would have to go out and search for things.
Nathan Lambert: You can tell it that you need to use search in this question.
Ross Taylor: Yeah, if you haven’t done a search, then you're probably doing it wrong. So yeah, in the long term, it gets washed out because there's nothing a neural network can't do compared to a human. But in the short term, there's still a lot of nooks and crannies that a model wouldn't quite cover / struggle on.
Nathan Lambert: Can you create a generative reward model by training off a bunch of rubric data? Probably?
Ross Taylor: Yeah, so verification benefits from thinking time. And I think most people are aware of this now, but it's more of a question of how you actually execute that. So a generative reward model for something like math and code - where it's like a 0 or 1 reward that you’re trying to figure out by thinking - is less interesting to me then questions where you really need to think from first principles on how to assign reward.
In general, the simplest way I think about it is: if you're moving to a world where you have long agentic traces, then your “reward model” just needs to answer a simple question, which is: “is the agent making progress towards its goal?” Right? But that's a very deep question.
So if it's a Pokemon eval, then maybe a model needs to use its knowledge of Pokemon to figure out if the agent in a trajectory has been caught in a loop, and whether it should be going towards Lavender town instead of this other way.
So these sorts of verification tasks benefit from thinking time, but the devil is in the detail. Because if you’re not careful, you’re just going to spend an inordinate amount of compute trying to get a reward.
Nathan Lambert: It feels like there will be a lot more we will learn there. It feels obviously salient. I’d describe it as verification changing the slope of inference time scaling. And that's really, really valuable if you're spending a lot on inference, but we don't really know how to do this. Like parallel compute is another factor that changes the shape of that curve.
I guess it's really all a slope of a scaling law or like an offset or something, but it's hard to say which things are true in terms of what we're hearing. That's probably what they're doing other than this rubric stuff. It's just a way to get RL pointed at more problems, which is not surprising.
Ross Taylor: Yeah, I think RubricMania is in full force right now. I mean, I think the longer term question, which has been posed in several places, is what happens when verification becomes fundamentally harder?
So I'm quite interested in the scientific discovery question. But in a field like biology, you need to do a physical experiment in order to verify. So it’s not just a question of running things on a cluster. And if you want to simulate the underlying thing, well then you’re bottlenecked by the quality of the simulation - and it turns out to be quite hard to simulate some physical processes!
Actually - in most of the sciences - I think this is the other point I’d make” which is that in ML, people overvalue the value of individual “thinking” in something like science. They think of Einstein and they think a lot less about the data generating mechanism, and what's the instrument.
There is no Kepler without the telescope. There is no progress in biology without X-ray crystallography. There's maybe new theories on dark matter in space without even better, newer telescopes.
I know this sounds like a weird say in the context of RL, but if you’re thinking about very hard things to solve in the real world, then you’re just going to be bottlenecked by the need to build a better instrument to get data. So it sounds like a digression, but I’m saying that - in the long-term - you’re going to hit these bottlenecks for verification. But in the short-term, we can still solve very interesting things like Millenium Prize problems, but that will probably take quite a while too!
Nathan Lambert: Yeah, I don't have anything particularly eloquent to say on the scientific discovery point. I guess what will happen is that RL is going to be in training and then you just sort of punt it off to the rest of post-training. So models need to be able to get really weird, but not weird in a way that they are numerically lost.
I've been reading a lot of reasoning traces these days, and the Qwen and DeepSeek reasoning traces really just seem numerically lost for a while, and then they eventually get the answer right. They say “Wait” a lot and then go into half English/half Chinese, and end up getting the answer right.
My point is that I don’t think in their current form, that these things are vehicles towards (scientific) discovery. There’s some kind of fundamental research needed to make the reasoning process more real.
Ross Taylor: My other bear case against reasoning models is the following argument - and this is mainly a devil’s advocate point, because I still fundamentally believe. Since World War 2, there are a lot more scientists in the world. But has progress kept up at the same rate? If anything, I would say that scientific progress has slowed.
Was there more progress in fundamental physics now or in the last century? And I know that is mainly because the low-hanging fruit is gone in many of these fields, but it could also be a bear case for AI because it hints that the bottleneck in science is the amount of intelligence on a problem, but maybe the speed of physical processes, or the ability to build better instruments for measuring, or the ability to get funding from governments to build bigger particle colliders…
I'm exaggerating the bear case because I think AGI mostly means autoating regular activities - law, finance and these kind of industries - and I think that’s a lot easier to do. But I’m attacking this mindset that says - now that we’ve solved reasoning - the takeoff is going to arrive in the next few years. From what I can see, that is very unrealistic.
Nathan Lambert: I'm very I'm bullish on AI being used and bearish on whatever superintelligence takes. I think we’re too compute constrained for a takeoff. I think AI is going to be very good for financialization and digitalization and seamlessly globalizing the Internet and making all information transfer and acquisition effectively free.
Ross Taylor: Yeah.
Nathan Lambert: Which is really good. And I think historically, the US is very well-positioned to capture this by making products that run on top of cheap AI models.
Ross Taylor: Yep.
Nathan Lambert: I wanted to ask you what AI you actually use. I don't know if I've ever asked you it's normally revealing.
Ross Taylor: Okay, so the base models we’re doing experiments on are mainly Qwen - Qwen 3, but also Qwen 2 because we know the kind of quirks of that model a bit more. A lot of people do that. Then we also do some distillation jobs, where we’re mainly using DeepSeek-R1. We did use Kimi recently, but we didn’t see massive benefits for the benchmarks we were looking at.
Then from a personal productivity perspective, Claude Code is very, very good. My main worry with Claude Code is that - I think there's a paper on this - but people confuse agents making you more productive versus preventing you from exerting mental effort. So sometimes I'll have a day with Claude code where I feel like I use very little mental effort - and it feels amazing - but I'm pretty sure I've done less work.
That will change because the models get better, but I'm trying to teach myself to be a bit careful because sometimes I need to stay in control.
Nathan Lambert: It does seem like an equilibrium. I'm happy with it. I don't want to have to grind out some plotting code. I'm just gonna watch some sports highlights and let it do it for me. That's fine…
Ross Taylor: Yeah. But in general, there is a lot of positive feedback from the community on Claude Code. It’s a very impressive product for me.
Nathan Lambert: What is the niche of your use case, or is it a bunch of things? Is there something you could endorse? Do you use it in math or code tasks? Do you use it in your startup’s codebase?
Ross Taylor: It tends to be better with brand new codebases. But I mostly use it for tasks which are quite horizontally scalable. So I'll have some basic specification where I'll provide it with some example code of mine, and then say “here's what a good implementation looks like”, but I need this modification or twist done. Sorry, I'm being very vague because I don't want to talk about specifics, but…
Nathan Lambert: Yeah.
Ross Taylor: It tends to be better for that. And, yeah, where it becomes really bad is when the file size becomes too long. Then the agent tends to struggle and get into these weird line search doom loops. So, yeah, there's a bit of work to do where you have to structure the codebase a bit for it to be efficient. But in general, it’s quite helpful.
Nathan Lambert: It's such a success that pretty much everybody that tries it is doing at least small code projects with it. I think maybe since ChatGPT, there hasn’t been this strong of a reaction.
Is this like the GPT 3.5 level? Like, Claude 4 is like GPT 3.5, the original ChatGPT, and then a couple iterations it’s gonna be incredible...
Ross Taylor: Yeah, I guess the people who really appreciate Claude Code are developers. Right? But it doesn't have the mass appeal of ChatGPT, which could generate poetry or whatever at the time, which was the killer mainstream use case at the time…it sounds crazy now.
Nathan Lambert: But I guess pay for Claude Code. People won't pay for ChatGPT (laughs)...
Ross Taylor: Exactly. So maybe it's a better business model…
But, yeah, I think that's a good question. I wouldn't say it's a ChatGPT moment, but I would say it's probably one of the most impactful products since ChatGPT. It’s not a ChatGPT moment because it hasn’t got mainstream appeal yet. And the question is: what does that agent look like? I'm still shocked that Apple hasn't done anything yet because, for me, that would be the killer thing. We'll see if they get that s**t together.
But, yeah, I'd imagine it would be some sort of on-device model. That would be my guess. We’ll see
Nathan Lambert: Yeah, that’s fun. Did you also wanna mention AlphaEvolve? I've been so burnt by Google's hypey projects - like their chip design and stuff.
Is this like the AlphaGo story, where if you have a really high performance simulator, that’s well matched to a task and you can scale RL - like many actors in parallel - then you can get high performance? I talked to Eugene Vinitsky recently, one of my friends from Berkeley. And they were at Apple and they did this really parallel RL for self driving simulator, which was really awesome.
Is AlphaEvolve somewhat away from that, but is in the same vein of extracting simulators?
Ross Taylor: I think AlphaEvolve is very cool. In my mind, it's very interesting because it feels like we are going full circle. In the 90s, the cool things which didn’t quite work were genetic algorithms and neural networks. And it feels we often see a new lease of life for several algorithms once the right context develops and other components get in place
So in the case of AlphaEvolve, you're exploiting the strong latent knowledge of a neural network, but then you also have a neurosymbolic element….don’t read too much into that, Gary Marcus… where you have a database where you store past programs. And having that prior in the form of past programs is a very good way to exploit the internal creativity of a language model as opposed to creating from scratch each time.
Nathan Lambert: How does AlphaEvolve actually do this? I think a lot of people are not going to know what it is doing. I don't think I have a good knowledge of it.
Ross Taylor: Say you have a kernel optimization task. For example, you’re making good kernels for common ML architectures. So you start with a reference implementation, and then in essence, it's a bit like in- context learning where you’re taking that implementation and saying “propose a change”, and then you benchmark it and get a score. And then you have a database where you store that program and its score.
And then when you sample a new round, you have an algorithm - it tends to be based on island based algorithms - where you sample in proportion to the score but you also wanna explore a bit. And that's your new prior. So you're iterating and evolving a program.
Nathan Lambert: And this is just handed off to the language model? What is the language model actually inferencing? Is it inferencing new programs?
Ross Taylor: Yes. So imagine you're constructing your prompt. You fetch a past implementation from your database and it goes in. It probably has the score as well saying “this implementation above got this result”. Then you ask the model to propose a new change.
I am oversimplifying, but this is the essence of the approach. You propose a new change, you write a program, get the score, store it in a database, and then go again.
So, basically: anything where you can pose a neat optimization task, this algorithm tends to work very well.
This is a broader debate now about how AlphaEvolve compares to RL approaches. First of all, I think they can be complementary, but…
Nathan Lambert: Maybe the language model is trained with RL, I bet?
Ross Taylor: Yes, that too. The interesting thing by the way is that the bulk of the AlphaEvolve approach was not using the strongest Gemini model - they used a weaker model with faster inference. So that’s an interesting tidbit which is sort of anti model-scaling pilled. There is a nice balance to be found there…
But yes: back to RL vs AlphaEvolve. I think this is part of a broader trend on how you use compute and whether the approach is parallel or sequential. The AlphaEvolve approach benefits from parallelisation, but they’re not going into deep long reasoning traces (sequential) just yet. But you could use both approaches.
Similarly, with RL you usually solve problems from scratch. But you could also think of ways you might want to exploit good priors in the context window. Benchmarks like KernelBench sort of do that anyway, but they don’t evolve the reference implementation like AlphaEvolve does.
So I think it's definitely something to watch. I think AlphaEvolve is underhyped, but we’ll see many more papers on this direction soon.
Nathan Lambert: It seems like a sign of things to come - figuring out parallel compute in the right way. It might be that the biggest model doesn’t necessarily benefit the most from a parallel compute setting.
Ross Taylor: Yeah.
Nathan Lambert: I mean, there's a lot of ways you could think about this. Like, the guess is a 100 times cheaper and half as good…
Ross Taylor: Yeah. So maybe this is a bullshitty philosophical point, but think about it this way. In the past 5,000 years, humans have made a lot of progress, but their brains fundamentally haven’t changed. What makes us smarter is that we’ve followed an invention curriculum, where the next invention builds on previous inventions.
So in the RL context, that raises the question: would you rather start from scratch each time, or would you use the best thing you have and successfully iterate that by standing on someone else’s shoulders?
So this is definitely something to watch in RL space. Instead of AlphaZero-ing things from scratch, how do we maintain existing implementations and iterate upon those?
This is also related to how we develop language models, and the discussion we had about Claude Code. You can imagine having an agentic model that is very good for starting from scratch, but you could also have a model that's very good at dealing with an existing code base. And the question is which is more valuable? And the answer is both. But then depending on how you actually use those models, you might end up preferring a different model.
So I am trying to put AlphaEvolve into a much bigger context here: and see it as a bigger trend about how we use compute, but also how a model might learn to improve on a problem.
Nathan Lambert: Yeah, that's fun. There's going to be a lot more things like AlphaEvolve - where people with particular domain expertise do the muddling and figure things out and more things will fall out. It is very remarkable that a zero order optimizer like a genetic algorithm, just using prompts for language models, can get anything useful out. That is a major win for language models being a fundamental unit of compute.
Ross Taylor: Yeah, absolutely. And a major win for LLMs and creativity, right? Because the meme is like “Oh, LLMs can't be creative”, and I’m always thinking, at a fundamental level, the softmax is quite an expressive operation…You’ll get creativity eventually. It's just a question of how quickly you can pick it out from what you sample.
So, I think AlphaEvolve is also proof of creativity. You found many new state-of-the-art implementations in AlphaEvolve - and will see more to come in upcoming papers.
Nathan Lambert: I would also guess there's people doing stuff like that that don't publish it. Or they've taken different models and hill climbed in their domain by setting up these weird loops.
I think this is a good place to end things. I’m kind of fading. Thanks for coming back. I’m doing a trip to London at some point. I don’t think we’ve ever met in person, but that’ll happen at some point!
I think we're I mean, I'm kind of fading, so I think it's good. Thanks for coming back. I'm doing trip to London at some point. I don't think we've never met in person, but that'll happen at some point.
Good to see you.
Ross Taylor: Yeah, good to see you Nathan. I'll see you in a bit!
Today, the White House released its AI Action Plan, the document we’ve been waiting for to understand how the new administration plans to achieve “global dominance in artificial intelligence (AI).” There’s a lot to unpack in this document, which you’ll be hearing a lot about from the entire AI ecosystem. This post covers one narrow piece of the puzzle — its limited comments on open models and AI research investment.
For some context, I was a co-author on the Ai2 official comment to the Office of Science and Technology Policy (OSTP) for the AI Action Plan and have had some private discussions with White House staff on the state of the AI ecosystem.
A focus of mine through this document is how the government can enable better fully open models to exist, rather than just more AI research in general, as we’re in a shrinking time window where if we don’t create better fully open models then the academic community could be left with a bunch of compute to do research on models that are not reflective of the frontier of performance and behavior. This is why I give myself ~18 months to finish The American DeepSeek Project.
Important context for this document is to consider what the federal government can actually do to make changes here. The executive branch has limited levers it can pull to disperse funding and make rules, but it sends important signaling to the rest of the government and private sector.
Overall, the White House AI Action Plan comes across very clearly that we should increase investment in open models, and for the right reasons.
This reflects a shift from previous federal policy, where the Biden executive order had little to say about open models other than them getting grouped into models needing pre-release testing if they were trained with more than 10^26 FLOPS (which led to substantial discussion on the general uselessness of compute thresholds as a policy intervention). Later, the National Telecommunications and Information Administration (NTIA) released a report from under the umbrella of the Biden Administration that was far more positive on open models, but much more limited in the scope of its ability for agenda setting.
This is formatted as comments in line with the full text on open models and related topics in the action plan. Let’s dive in, any emphasis in italics is mine.
Encourage Open-Source and Open-Weight AI
Open-source and open-weight AI models are made freely available by developers for anyone in the world to download and modify. Models distributed this way have unique value for innovation because startups can use them flexibly without being dependent on a closed model provider. They also benefit commercial and government adoption of AI because many businesses and governments have sensitive data that they cannot send to closed model vendors. And they are essential for academic research, which often relies on access to the weights and training data of a model to perform scientifically rigorous experiments.
This covers three things we’re seeing play out with open models and is quite sensible as an introduction:
* Startups use open models to a large extent because pretraining themselves is expensive and modifying the model layer of the stack can provide a lot of flexibility with low serving costs. Today, most of this happens on Qwen at startups, where larger companies are more hesitant to adopt Chinese models.
* Open model deployments are slowly building up around sensitive data domains such as health care.
* Researchers need strong and transparent models to perform valuable research. This is the one I’m most interested in, as it is the one with the highest long-term impact by determining the fundamental pace of progress in the research community.
We need to ensure America has leading open models founded on American values. Open-source and open-weight models could become global standards in some areas of business and in academic research worldwide. For that reason, they also have geostrategic value. While the decision of whether and how to release an open or closed model is fundamentally up to the developer, the Federal government should create a supportive environment for open models.
The emphasized section is entirely the motivation behind ongoing efforts for The American DeepSeek Project. The interplay between the three groups above is inherently geopolitical, where Chinese model providers are actively trying to develop mindshare with Western developers and release model suites that offer great tools for research (e.g. Qwen).
The document is highlighting why fewer open models exist right now from leading Western AI companies, simply “the decision of whether and how to release an open or closed model is fundamentally up to the developer” — this means that the government itself can mostly just stay out of the way of leading labs releasing models if we think the artifacts will come from the likes of Anthropic, OpenAI, Google, etc. The other side of this is that we need to invest in building organizations around releasing strong open models for certain use cases that do not have economic conflicts or different foci.
Onto the policy steps.
Recommended Policy Actions
* Ensure access to large-scale computing power for startups and academics by improving the financial market for compute. Currently, a company seeking to use large-scale compute must often sign long-term contracts with hyperscalers—far beyond the budgetary reach of most academics and many startups. America has solved this problem before with other goods through financial markets, such as spot and forward markets for commodities. Through collaboration with industry, the National Institute of Standards and Technology (NIST) at the Department of Commerce (DOC), the Office of Science and Technology Policy (OSTP), and the National Science Foundation’s (NSF) National AI Research Resource (NAIRR) pilot, the Federal government can accelerate the maturation of a healthy financial market for compute.
The sort of issue the White House is alluding to here is that if you want to have 1000 GPUs as a startup or research laboratory you often need to sign a 2-3 year commitment in order to get low prices. Market prices for on-demand GPUs tend to be higher. The goal here is to make it possible for people to get the GPU chunks they need through financial incentives.
We’ve already seen a partial step for this in the recent budget bill, where AI training costs now can be classified as R&D expenses, but this largely helps big companies. Actions here that are even more beneficial for small groups releasing open weight or open-source models would be great to see.
One of the biggest problems I see for research funding is going to be the challenge of getting concentrated compute into the hands of researchers, so I hope the administration follows through here for compute density in places. A big pool of compute spread across the entire academic ecosystem means too little compute for models to get trained at any one location. It reads as if the OSTP understands this and has provided suitable guidance.
Interconnects is a reader-supported publication. Consider becoming a subscriber.
* Partner with leading technology companies to increase the research community’s access to world-class private sector computing, models, data, and software resources as part of the NAIRR pilot.
* Build the foundations for a lean and sustainable NAIRR operations capability that can connect an increasing number of researchers and educators across the country to critical AI resources.
This is simple and to my knowledge has largely been under way. NAIRR provided a variety of resources to many academic parties, such as API credits, data, and compute access, so it should be expanded upon. I wrote an entire piece on saving the NAIRR last November when its funding future was unclear (and needed Congressional action).
This is the balance to what I was talking about above on model training. It provides smaller resource chunks to many players, which is crucial, but doesn’t address the problem of building great open models.
* Continue to foster the next generation of AI breakthroughs by publishing a new National AI Research and Development (R&D) Strategic Plan, led by OSTP, to guide Federal AI research investments.
This seems like a nod to a logical next step.
Where the overall picture of research funding in the U.S. has been completely dire, the priority in AI research has already been expressed through AI being the only area of NSF grant areas without major cuts. There is likely to be many other direct effects of this, but it is out of scope of the article.
More exact numbers can be found in the NSF 2026 proposed budget, where AI is an outlier as one of the only topics with a positive net change from 2024 or 2025.
* Led by DOC through the National Telecommunications and Information Administration (NTIA), convene stakeholders to help drive adoption of open-source and open-weight models by small and medium-sized businesses.
This is a more unexpected line item, but a welcome one. It’ll be harder to implement, but if it works it’ll do a lot of good for building momentum around open model investment. A large part of why few open models exist in the U.S. is just because there’s not a lot of business value from releasing them. A big story of 2025 has been how open models are closing the gap in capabilities, or at least crossing important ability thresholds, which could start to change this equilibrium.
That’s it for the core section on open models! It’s right to the point.
There are a couple related sections I wanted to point you to, which largely complement the above or show how it is hard for a document like this to acknowledge things like “our R&D ecosystem is being outcompeted by Chinese models.”
First, more on AI research itself.
Advance the Science of AI
Just as LLMs and generative AI systems represented a paradigm shift in the science of AI, future breakthroughs may similarly transform what is possible with AI. It is imperative that the United States remain the leading pioneer of such breakthroughs, and this begins with strategic, targeted investment in the most promising paths at the frontier.
Recommended Policy Actions
* Prioritize investment in theoretical, computational, and experimental research to preserve America’s leadership in discovering new and transformative paradigms that advance the capabilities of AI, reflecting this priority in the forthcoming National AI R&D Strategic Plan.
Something in my mind that is very missing from this document is a comment on immigration. If we want the U.S. to be a leader in AI research we need to prioritize fixing the immigration ecosystem as soon as possible. Leading AI conferences can no longer be located solely in the U.S. because too many authors cannot get a travel visa in time to attend the conference, let alone the other issues on hiring or funding at academic institutions.
This section on the Science of AI reads very similar to the section on open models.
And the only mentions of China, which is related as the party pushing open models the furthest today:
Counter Chinese Influence in International Governance Bodies
A large number of international bodies, including the United Nations, the Organisation for Economic Co-operation and Development, G7, G20, International Telecommunication Union, Internet Corporation for Assigned Names and Numbers, and others have proposed AI governance frameworks and AI development strategies. The United States supports likeminded nations working together to encourage the development of AI in line with our shared values. But too many of these efforts have advocated for burdensome regulations, vague “codes of conduct” that promote cultural agendas that do not align with American values, or have been influenced by Chinese companies attempting to shape standards for facial recognition and surveillance.
Recommended Policy Actions
* Led by DOS and DOC, leverage the U.S. position in international diplomatic and standard-setting bodies to vigorously advocate for international AI governance approaches that promote innovation, reflect American values, and counter authoritarian influence.
and a quick comment on Chinese talking points in the section “Ensure that Frontier AI Protects Free Speech and American Values”:
* Led by DOC through NIST’s Center for AI Standards and Innovation (CAISI), conduct research and, as appropriate, publish evaluations of frontier models from the People’s Republic of China for alignment with Chinese Communist Party talking points and censorship.
This reads as there being a low probability that we see any immediate executive action trying to ban the likes of Qwen or DeepSeek, which is good for the time being. The evaluation of Chinese and American values is a slippery slope in some ways, as it quickly will become enmeshed in the idea of “woke AI,” but in the meantime it is likely to be a major talking point with respect to the open models we’re seeing from Chinese companies, which do often parrot very simple talking points reflective of “Chinese socialist values.”
We need our ecosystem to compete on merits of the technology being better at useful tasks if we want to lead in the long-term technological arc, rather than political games. That’s my number one focus over the next couple of years and why I reiterate the need for open models for fundamental AI research and innovation. The biggest beneficiaries of this sort of innovation have historically been the biggest American technology companies, who now should do their part to support them existing — with some government encouragement.
Let me know if I missed anything, as this was a quick pass to make sure I read the details and connected the recent dots.
https://www.interconnects.ai/p/kimi-k2-and-when-deepseek-moments
The DeepSeek R1 release earlier this year was more of a prequel than a one-off fluke in the trajectory of AI. Last week, a Chinese startup named Moonshot AI dropped Kimi K2, an open model that is permissively licensed and competitive with leading frontier models in the U.S. If you're interested in the geopolitics of AI and the rapid dissemination of the technology, this is going to represent another "DeepSeek moment" where much of the Western world — even those who consider themselves up-to-date with happenings of AI — need to change their expectations for the coming years.
In summary, Kimi K2 shows us that:
* HighFlyer, the organization that built DeepSeek, is far from a uniquely capable AI laboratory in China,
* China is continuing to approach (or reached) the absolute frontier of modeling performance, and
* The West is falling even further behind on open models.
Kimi K2, described as an "Open-Source Agentic Model" is a sparse mixture of experts (MoE) model with 1T total parameters (~1.5x DeepSeek V3/R1's 671B) and 32B active parameters (similar to DeepSeek V3/R1's 37B). It is a "non-thinking" model with leading performance numbers in coding and related agentic tasks (earning it many comparisons to Claude 3.5 Sonnet), which means it doesn't generate a long reasoning chain before answering, but it was still trained extensively with reinforcement learning. It clearly outperforms DeepSeek V3 on a variety of benchmarks, including SWE-Bench, LiveCodeBench, AIME, or GPQA, and comes with a base model released as well. It is the new best-available open model by a clear margin.
These facts with the points above all have useful parallels for what comes next:
* Controlling who can train cutting edge models is extremely difficult. More organizations will join this list of OpenAI, Anthropic, Google, Meta, xAI, Qwen, DeepSeek, Moonshot AI, etc. Where there is a concentration of talent and sufficient compute, excellent models are very possible. This is easier to do somewhere such as China or Europe where there is existing talent, but is not restricted to these localities.
* Kimi K2 was trained on 15.5T tokens and has a very similar number of active parameters as DeepSeek V3/R1, which was trained on 14.8T tokens. Better models are being trained without substantial increases in compute — these are referred to as a mix of "algorithmic gains" or "efficiency gains" in training. Compute restrictions will certainly slow this pace of progress on Chinese companies, but they are clearly not a binary on/off bottleneck on training.
* The gap between the leading open models from the Western research labs versus their Chinese counterparts is only increasing in magnitude. The best open model from an American company is, maybe, Llama-4-Maverick? Three Chinese organizations have released more useful models with more permissive licenses: DeepSeek, Moonshot AI, and Qwen. This comes at the same time that new inference-heavy products are coming online that'll benefit from the potential of cheaper, lower margin hosting options on open models relative to API counterparts (which tend to have high profit margins).
Kimi K2 is set up for a much slower style "DeepSeek Moment" than the DeepSeek R1 model that came out in January of this year because it lacks two culturally salient factors:
* DeepSeek R1 was revelatory because it was the first model to expose the reasoning trace to the users, causing massive adoption outside of the technical AI community, and
* The broader public is already aware that training leading AI models is actually very low cost once the technical expertise is built up (recall the DeepSeek V3 $5M training cost number), i.e. the final training run is cheap, so there should be a smaller reaction to similar cheap training cost numbers in the Kimi K2 report coming soon.
Still, as more noise is created around the K2 release (Moonshot releases a technical report soon), this could evolve very rapidly. We've already seen quick experiments spin up slotting it into the Claude Code application (because Kimi's API is Claude-compatible) and K2 topping many nice "vibe tests" or creativity benchmarks. There are also tons of fun technical details that I don't have time to go into — from using a relatively unproven optimizer Muon and scaling up the self-rewarding LLM-as-a-judge pipeline in post-training. A fun tidbit to show how much this matters relative to the noisy Grok 4 release last week is that Kimi K2 has already surpassed Grok 4 in API usage on the popular OpenRouter platform.
Later in the day on the 11th, following the K2 release, OpenAI CEO Sam Altman shared the following message regarding OpenAI's forthcoming open model (which I previously shared more optimistic thoughts on here) :
we planned to launch our open-weight model next week.
we are delaying it; we need time to run additional safety tests and review high-risk areas. we are not yet sure how long it will take us.
while we trust the community will build great things with this model, once weights are out, they can’t be pulled back. this is new for us and we want to get it right.
sorry to be the bearer of bad news; we are working super hard!
Many attributed this as a reactive move by OpenAI to get out from the shadow of Kimi K2's wonderful release and another DeepSeek media cycle.
Even though someone at OpenAI shared with me that the rumor that Kimi caused the delay for their open model is very likely not true, this is what being on the back foot looks like. When you're on the back foot, narratives like this are impossible to control.
We need leaders at the closed AI laboratories in the U.S. to rethink some of the long-term dynamics they're battling with R&D adoption. We need to mobilize funding for great, open science projects in the U.S. and Europe. Until then, this is what losing looks like if you want The West to be the long-term foundation of AI research and development. Kimi K2 has shown us that one "DeepSeek Moment" wasn't enough for us to make the changes we need, and hopefully we don't need a third.
https://www.interconnects.ai/p/the-american-deepseek-project
While America has the best AI models in Gemini, Claude, o3, etc. and the best infrastructure with Nvidia it’s rapidly losing its influence over the future directions of AI that unfold in the open-source and academic communities. Chinese organizations are releasing the most notable open models and datasets across all modalities, from text to robotics or video, and at the same time it’s common for researchers worldwide to read far more new research papers from Chinese organizations rather than their Western counterparts.
This balance of power has been shifting rapidly in the last 12 months and reflects shifting, structural advantages that Chinese companies have with open-source AI — China has more AI researchers, data, and an open-source default.
On the other hand, America’s open technological champions for AI, like Meta, are “reconsidering their open approach” after yet another expensive re-org and the political environment is dramatically reducing the interest of the world’s best scientists in coming to our country.
It’s famous lore of the AI industry that much of the flourishing of progress around ChatGPT is downstream from Google Research’s, and the industry’s writ-large, practice of openly sharing the science of AI until approximately 2022. Stopping this practice, and the resulting power shifts mean it will be likely that the next “Transformer”-style breakthrough will be built on or related to Chinese AI models, AI chips, ideas, or companies. Countless Chinese individuals are some of the best people I’ve worked with, both at a technical and personal level, but this direction for the ecosystem points to AI models being less accountable, auditable, and trustworthy due to inevitable ties to the Chinese Government.
The goal for my next few years of work is what I’m calling The American DeepSeek Project — a fully open-source model at the scale and performance of current (publicly available) frontier models, within 2 years. A fully open model, as opposed to just an “open weights” model, comes with data, training code, logs, and decision making — on top of the weights to run inference — in order to distribute the knowledge and access for how to train AI models fully.
This project serves two goals, where balancing the scales with the pace of the Chinese ecosystem is only one piece:
* Reclaim the AI research default home being on top of American (or Western) technologies and tools, and
* Reduce the risk that the only viable AI ecosystem for cutting edge products in built atop of proprietary, closed, for-profit AI models.
More people should be focused on this happening. A lot of people talk about how nice it would be to have “open-source AGI for all,” but very few people are investing in making it reality. With the right focus, I estimate this will take ~$100M-500M over the next two years.
Within the context of recent trends, this is a future that has a diminishing, minute probability. I want to do this at Ai2, but it takes far more than just us to make it happen. We need advocates, peers, advisors, and compute.
The time to do this is now, if we wait then the future will be in the balance of extremely powerful, closed American models counterbalancing a sea of strong, ubiquitous, open Chinese models. This is a world where the most available models are the hardest to trust. The West historically has better systems to create AI models that are trustworthy and fair across society. Consider how:
* Practically speaking, there will never be proof that Chinese models cannot leave vulnerabilities in code or execute tools in malicious ways, even though it’s very unlikely in the near future.
* Chinese companies will not engage as completely in the U.S. legal system on topics from fair use or non-consensual deepfakes.
* Chinese models will over time shift to support a competitive software ecosystem that weakens many of America and the West’s strongest companies due to in-place compute restrictions.
Many of these practical problems cannot be fixed by simply fine-tuning the model, such as Perplexity’s R1-1776 model. These are deep, structural realities that can only be avoided with different incentives and pretrained models.
My goal is to make a fully open-source model at the scale of DeepSeek V3/R1 in the next two years. I’ve been starting to champion this vision in multiple places that summarizes the next frontier for performance on open-source language models, so I needed this document to pin it down.
I use scale and not performance as a reference point for the goal because the models we’re collectively using as consumers of the AI industry haven’t really been getting much bigger. This “frontier scale” is a ballpark for where you’ve crossed into a very serious model, and, by the time a few years has gone by, the efficiency gains that would’ve accumulated by then will mean this model will far outperform DeepSeek V3. The leading models used for synthetic data (and maybe served to some users) will continue to get bigger, but not as quickly as capabilities will grow and new types of agents will emerge.
Interconnects is a reader-supported publication. Consider becoming a subscriber.
The terminology “American DeepSeek” is stretching words in order to be identifiable to a broad public. It combines the need for true American values with a breakthrough open release that marks a new milestone in capabilities.
DeepSeek is known for many things to the general public — training cheap frontier models, bringing reasoning models to consumers, and largely being the face of Chinese AI efforts. Since ChatGPT, DeepSeek is the first organization to release an open, permissively licensed AI model at the frontier of performance. This was a major milestone and why 2025 has been a transformative year in the perception of feasibility for open models generally. The name DeepSeek will forever be known in AI lore for it.
At the same time, what will count as a “DeepSeek moment” is changing. The new directions for where AI is heading is more in line with agents that use models a lot (sometimes even smaller models) rather than relying on scaling performance of single model generations.
This changes what it’ll mean for models to be “at the frontier.” More releases will look like Claude 4 and be about usability, where the benchmarks that people are hillclimbing on represent new types of capabilities or outlandish, harder than human expert tasks. For the suite of tasks that were core for the current generation of models: MATH, GPQA, SWE-Bench Verified, etc., solving them represents a challenging, but reasonable, baseline for human performance.
The next major milestone will be when fully open-source models reach this performance threshold. With fully open-source models at this level, “anyone” can specialize the model to their task and the possibility of an open ecosystem that runs efficiently on a single architecture can coalesce. This doesn’t mean releasing the best AI models of 2027 with complete openness — just that we should, come 2027, have fully open models of 2025’s capabilities in order to enable new types of companies and research.
The efficiencies of open-source software style development are dramatically stronger for agentic systems than models. Models are singular entities built with expensive resources and incredible focus. Agents are systems that can use many models off the shelf and route requests depending on what’s needed.
This agentic era is the opportunity open models have needed, but we need to clear much stronger performance thresholds before the open counterparts are viable. We have companies like OpenAI and Google launching Claude Code competitors that pretty much flop. Imagine what this would look like with open models today? Not good.
For this reason, we have finite time to get there. Surely, eventually this level of models will exist, but if we want a new type of ecosystem to form we need to build the raw resources while developers and new companies are getting started. We need people willing to take the risk on something different while there is still potential for it to be comparable across performance trade-offs.
Today, the best fully open language models are catching up to the levels of the original GPT-4. This is a major step from GPT-3 levels. The required step I’m shooting for is reaching the modern GPT-4 type models, the likes of recent Sonnet, DeepSeek V3, or Gemini Pro. It’s a big step, but a transformative one in terms of what the models can do.
Of course, some of this still works with open weight models and not just fully open models, but to date we have not had good success with having open weight models that can fully be trusted. The best American models are plagued by the Llama license (and rumors that future versions will be discontinued). At the same time, Chinese models aren’t trusted because the models are being integrated directly with more complex tools that muddy the water with a weak security reputation, and European models are largely off the map.
If we want models we can trust, we need something that’s a bit different. If the models all converge on a certain capability level, and the differentiation is on integration and finetuning to specific skills, this is something the open community can do.
In many ways, obtaining this goal is a quintessentially American volition. In the face of a technology that is poised to bring such extreme financial, and by proxy literal, power to a few companies, opening AI is one of the only things we can do to reduce it. Technology proceeds in a one-way direction — for a variety of geopolitical and capitalistic reasons it is impractical to pause AI development to “do AI another way” — the best we can do is chart a path that makes this future better.
Along the same vein, if AGI already exists and something closer to ASI is coming, it will be intertwined with countless details of billions of people’s lives in a matter of just years. Something so indispensable to our lives in work, play, entertainment, and relationships is a closer analog to electricity than other traditional technology products that one can opt into. Such technology should be available for all to benefit from.
We need new systems to mitigate misuse, but it shouldn’t be solely up to corporations to control this. Safety by isolating technology to a select few is something we’re in the later stages of with nuclear weapons, and AI progress is far harder to monitor. Robustness to AI can only come from designing systems that expect it to be pervasive — not that it is an easy task.
Realistically, all of this is fighting gravity. The corporations will win, but we can control to what extent. We can control how good the other options are. The open options.
The call to action here is simple — consider how you can slightly shift your decision making to make The American DeepSeek more likely. This approach succeeds just as much by having one model at the end of it, as it does by having the community form better habits and norms around the way AI models are conceived, built, shared, and used.
https://www.interconnects.ai/p/summertime-outlook-o3s-novelty-coming
Summer is always a slow time for the tech industry. OpenAI seems fully in line with this, with their open model “[taking] a little more time” and GPT-5 seemingly always delayed a bit more. These will obviously be major news items, but I’m not sure we see them until August.
I’m going to take this brief reprieve in the bombardment of AI releases to reflect on where we’ve been and where we’re going. Here’s what you should know.
1. o3 as a technical breakthrough beyond scaling
The default story around OpenAI’s o3 model is that they “scaled compute for reinforcement learning training,” which caused some weird, entirely new over-optimization issues. This is true, and the plot from the livestream of the release still represents a certain type of breakthrough — namely scaling up data and training infrastructure for reinforcement learning with verifiable rewards (RLVR).
The part of o3 that isn’t talked about enough is how different its search feels. For a normal query, o3 can look at 10s of websites. The best description I’ve heard of its relentlessness en route to finding a niche piece of information is akin to a “trained hunting dog on the scent.” o3 just feels like a model that can find information in a totally different way than anything out there.
The kicker with this is that we’re multiple months out from its release in April of 2025 and no other leading lab has a model remotely like it. In a world where releases between labs, especially OpenAI and Google, seem totally mirrored, this relentless search capability in o3 still stands out to me.
The core question is when will another laboratory release a model that feels qualitatively similar? If this trend goes on through the end of the summer it’ll be a confirmation that OpenAI had some technical breakthrough to increase the reliability of search and other tool-use within reasoning models.
For a contrast, consider basic questions we are facing in the open and academic community on how to build a model inspired by o3 (so something more like a GPT-4o or Claude 4 in its actual search abilities):
* Finding RL data where the model is incentivized to search is critical. It’s easy in an RL experiment to tell the model to try searching in the system prompt, but as training goes on if the tool isn’t useful the model will learn to stop using it (very rapidly). It is likely that OpenAI, particularly combined with lessons from Deep Research’s RL training (which, I know, is built on o3), has serious expertise here. A research paper showing a DeepSeek R1 style scaled RL training along with consistent tool use rates across certain data subsets will be very impressive to me.
* The underlying search index is crucial. OpenAI’s models operate on a Bing backend. Anthropic uses Brave’s API and it struggles for it (lots of SEO spam). Spinning up an academic baseline with these APIs is a moderate additive cost on top compute.Once solid open baselines exist, we could do fun science such as studying which model can generalize to unseen data-stores best — a crucial feature for spinning up a model on local sensitive data, e.g. in healthcare or banking.
If you haven’t been using o3 for search, you really should give it a go.
Interconnects is a reader-supported publication. Consider becoming a subscriber.
2. Progress on agents will be higher variance than modeling was, but often still extremely rapid
Claude Code’s product market fit, especially with Claude 4, is phenomenal. It’s the full package for a product — works quite often and well, a beautiful UX that mirrors the domain, good timing, etc. It’s just a joy to use.
With this context, I really have been looking for more ways to write about it. The problem with Claude Code, and other coding agents such as Codex and Jules, is that I’m not in the core audience. I’m not regularly building in complex codebases — I’m more of a research manager and fixer across the organization than someone that is building in one repository all the time — so, I don’t have practical guides on how to get the most out of Claude Code or a deep connection with it that can help you “feel the AGI.”
What I do know about is models and systems, and there are some very basic facts of frontier models that make the trajectory for the capabilities of these agents quite optimistic.
The new part of LLM-based agents is that they involve many model calls, sometimes with multiple models and multiple prompt configurations. Previously, the models everyone was using in chat windows were designed to make progress on linear tasks and return that to the user — there wasn’t a complex memory or environment to manage.
Adding a real environment for the models has made it so the models need to do more things and often a wider breadth of tasks. When building these agentic systems, there are two types of bottlenecks:
* The models cannot solve any of the task we hope to use the agent for, and
* The models fail at small components of the task that we are deploying.
For agents that have initial traction, such as Claude Code and Deep Research, many of the problems are in the second class. How these fixes are made is that labs notice repeated, odd failures among real world use-cases. This can look like a 50% reliability rate on some long-tail mundane task. In this case it is often easy for the lab to make new data, include it in the next post-training run for their models, and up that sub-task reliability to almost 99%. As labs are making most of their gains in post-training today, rather than big pretraining runs, the time for that change to get integrated is well shorter than recent years.
The kicker for this is how it all fits together. Many complex tasks can be bottlenecked by some weird, small failures. In this case, we can have small changes to models that make agents like Claude Code feel way more reliable, even though the peak performance of the model hasn’t changed much. The same goes for Deep Research.
With this, I expect these agents we’re already using to improve randomly and in big leaps.
What I’m unsure of is when new agent platforms will be built. Some of this is a product problem and some of it is a peak performance problem. New agentic platforms that feel like they have product-market fit will be somewhat random, but those that have a fit already can improve like we’re used to frontier models getting way better.
This is a different path for the industry and will take a different form of messaging than we’re used to. More releases are going to look like Anthropic’s Claude 4, where the benchmark gains are minor and the real world gains are a big step. There are plenty of more implications for policy, evaluation, and transparency that come with this. It is going to take much more nuance to understand if the pace of progress is continuing, especially as critics of AI are going to seize the opportunity of evaluations flatlining to say that AI is no longer working.
Much like o3, you should play with Claude Code even if you don’t code a lot. It can make fun demos and standalone websites in no time. It’s miles ahead in its approachability compared to the fully-autonomous agents like Codex (at least for the time being).
3. Scaling parameters is going to go very slow for consumer models
The models that leading AI labs have been releasing in 2025 have mostly stopped getting bigger in total parameters. Take Claude 4, the API prices are the same as Claude 3.5 (and its minor versions). OpenAI only half released GPT-4.5. Gemini hasn’t released its Ultra variant. There are more models that are private to these laboratories that are certainly much bigger.
The nuanced part of this is that many of these models likely could be getting slightly smaller, e.g. Claude 4 Sonnet could be slightly smaller than Claude 3.5 Sonnet, due to efficiency gains at pretraining. That sort of marginal technical advancement is a big deal on price and inference speed, especially in the long-run, but not the central point I’m making.
The point is how GPT-5 is going to be bigger mostly through inference-time scaling and less through just “one bigger model.” For years we were told the narrative that the lab with the biggest training cluster was going to win because they have an advantage with scaling. That was the story behind xAI’s mega-cluster that Elon built. Now, the biggest cluster just is an advantage in overall research pace.
Scaling, at least in terms of what users need, has largely fizzled out. Labs may come back to it later as they find super hard problems that users need to solve, but where GPT 4.5 cost about 100x the compute of GPT-4 to train, it is only slightly better on normal user metrics.
What we see now is a mass efficiency march along the model sizes that people love. The industry has a few standards, from
* Tiny models like Gemini Flash Lite or GPT 4.1 Nano,
* Small models like Gemini Flash and Claude Haiku,
* Standard models like GPT-4o and Gemini Pro, and
* Big models like Claude Opus and Gemini Ultra.
These models come with somewhat predictable price-points (we know Gemini is way cheaper than the industry standard), latencies, and capability levels. Standards like this are important as industries mature!
Over time, efficiency gains will make new standards emerge. The first thing we’ll see is more mass availability of the likes of Gemini Ultra and GPT-4.5 (maybe in the GPT-5 release), but what comes after that isn’t on the radar at all. Now, scaling to new size tiers is only possible “every few years” or maybe not at all, if monetization of AI doesn’t go as well as many hope.
Scaling as a product differentiator died in 2024. That doesn’t mean pretraining as a science isn’t crucial. The recent Gemini 2.5 report made that pretty clear:
The Gemini 2.5 model series makes considerable progress in enhancing large-scale training stability, signal propagation and optimization dynamics, resulting in a considerable boost in performance straight out of pre-training compared to previous Gemini models.
From the publisher's feed

542 Listeners

1,089 Listeners

288 Listeners

203 Listeners

205 Listeners

315 Listeners

98 Listeners

565 Listeners

141 Listeners

102 Listeners

222 Listeners

145 Listeners

455 Listeners

30 Listeners

39 Listeners