
Sign up to save your podcasts
Or


By Latent.Space
4.6
9292 ratings
The podcast currently has 249 episodes available.
The most played episodes among Podcast App listeners.

At 1:09:00 we talk about the rise of AI x Finance, and AIE NYC is one month away - our hotel block is 97% sold out, get tix & travel ASAP - we will announce speakers from Bridgewater, Ramp, Coatue, Mastercard, Vanguard, Coinbase, Blackrock, Fidelity, Point72, Capital One, JPMC, Wells Fargo, Bloomberg, A24 (yes the movie studio) Labs, Two Sigma, Apollo Global, and more soon! From helping pioneer core ideas in NLP to now building AI systems that can automate AI research itself, Richard Socher is betting that the next major step in AI is recursive self-improvement. He is the founder of You.com, AIX Ventures, and now Recursive, which has assembled some of the best open-endedness (& self improving agent) researchers in the world and raised a $4.65B seed round. In this episode, Richard joins Latent Space to unpack his vision for the “Eureka Machine”: a superintelligence that can improve the process of invention itself, accelerate AI research, and eventually tackle major problems across science, energy, materials, biology, and more. You can get his book “The Eureka Machine” here! We go deep on Recursive’s early results, including an AI research system that Richard says outperformed humans and their agents on optimization tasks in less than two days, as well as work on NVIDIA GPU kernels where the system discovered improvements without relying on a team of CUDA experts. Richard also explains why he thinks AI research that currently takes thousands of people and years could eventually be compressed into weeks. These results are summarized in his 20 minute AIE keynote, where we also discuss his 10 dimensions of intelligence: We also explore the harder questions around increasingly capable AI: reward hacking, whether Anthropic-style constitutions actually work, AI regulation and proposals to “pace” frontier development, open-source models as geopolitical soft power, whether today’s LLM paradigm is enough, and what happens if AI systems eventually begin choosing their own goals. Richard reflects on the rejected research that helped inspire Alec Radford’s GPT, open-endedness, the AI Economist, simulations of entire economies, and his framework for thinking about the upper bounds of intelligence itself. We discuss: * The Eureka Machine and Richard’s vision for an AI that can automate invention * Why Richard is optimistic about superintelligence for science and technology * Why AI hard-takeoff scenarios may underestimate physical and economic constraints * The risks of regulating intelligence itself instead of specific AI applications * Reward hacking and why increasingly intelligent AI makes objective design harder * Richard’s critique of Anthropic’s constitution and constitutional AI * Alignment vs. personalization and whose values an AI should follow * Why open-source AI matters for resilience, competition, and geopolitical soft power * Why Richard left You.com’s frontier-model work to start Recursive * Recursive self-improvement and automating the process of AI research * Whether today’s LLM paradigm is enough — and why Richard is less bullish on world models * DecaNLP, early prompt-based generalization, and the research that influenced GPT * Why rejected research can shape entire technological timelines * Open-endedness, evolutionary approaches, and rainbow teaming * What happens if AI systems begin setting their own goals * Why simple objectives like profit maximization can produce dangerous reward hacks * Recursive’s long-term plan to apply self-improving AI to science * The compute, hardware, and economic constraints on AI takeoff * Recursive’s early NanoChat, NanoGPT, and GPU kernel optimization results * Why automating AI research could reduce years of work to weeks * Reward engineering and what makes auto-research systems actually work * The AI Economist and using simulations to test economic policy * Whether LLMs can realistically simulate people and entire economies * Benchmark bugs and evaluation harnesses and the difficulty of measuring AI progress * Recursive’s near-term focus on AI for AI research * Harness optimization, sandboxing, and web search as core agent infrastructure * You.com and the search stack for AI agents * AI in finance, backtesting, and data leakage * Richard’s three fundamental components and ten “spaces” of intelligence * The theoretical upper bounds of vision, communication, knowledge, and computation * Creative intelligence, metacognition, and AI-generated goals * Survival and replication and why AI does not necessarily need to fear being turned off * High agency and ambitious goals and Richard’s advice for people building with AI Richard Socher * X: https://x.com/RichardSocher * LinkedIn: https://www.linkedin.com/in/richardsocher/ Timestamps 00:00:00 The Eureka Machine and Superintelligence 00:02:23 AI Optimism, Slow Takeoff, and Regulation 00:07:56 AI Safety, Reward Hacking, and Anthropic’s Constitution 00:11:49 Alignment, Personalization, and Open Source AI 00:15:46 Why Richard Started Recursive 00:20:03 Recursive Self-Improvement and the Founding Team 00:22:55 Are Today’s LLMs Enough? 00:29:03 DecaNLP, GPT, and the Rejected Idea Ahead of Its Time 00:34:38 Open-Endedness and Evolutionary AI 00:36:38 What Happens When AI Chooses Its Own Goals? 00:41:16 Superintelligence for Science 00:42:40 GPUs, Compute, and the Limits of AI Takeoff 00:45:07 Recursive’s Results: AI Beating Humans and Their Agents 00:49:14 Reward Engineering and Auto Research 00:53:12 The AI Economist and Simulating Entire Economies 00:58:07 LLM Simulations, Personas, and Mode Collapse 01:03:38 Recursive’s Roadmap, Agents, Search, and Finance 01:09:13 The Upper Bounds and Spaces of Intelligence 01:30:21 Goals, High Agency, and Advice for Builders Transcript Introduction: Richard Socher and the Eureka Machine Swyx [00:00:00]: We’re here in a studio with Vibhu and myself and Richard Socher. Welcome. Richard Socher [00:00:06]: Thanks for having me. Swyx [00:00:07]: We just talked about the Eureka Machine, or we just released a talk, at AI Engineer about the Eureka Machine. Is it — you said it’s your life’s goal. What is the Eureka Machine? Richard Socher [00:00:16]: The Eureka Machine is the ultimate invention that will afterwards invent most everything for humanity. It’s essentially a superintelligence that can be given any goal, any environment, reward, and then it will try its best to achieve those goals to create the kinds of inventions that humanity would hopefully ask it for. Swyx [00:00:45]: Yeah, I think we have the book pulled up here that you’ve written. Richard Socher [00:00:50]: That’s right, yeah. I finished it last year, a little bit before we started Recursive, and now we’re gonna try to build parts of that. Swyx [00:00:57]: You finished it last year. It’s July. What takes so long? Richard Socher [00:01:01]: Oh, man, books. Books are incredibly slow. Richard Socher [00:01:04]: It’s ridiculous. That whole industry is just unfathomably slow. Richard Socher [00:01:07]: So a lot of the ideas have been out there for a while, but yeah, I’m really glad it’s finally coming out in September this year. Swyx [00:01:14]: We might have AGI by then. Like, we don’t know. Vibhu [00:01:18]: Any key takeaway that you’re most excited to put in here? Techno-Optimism, AI Upside, and Slow Takeoff Richard Socher [00:01:21]: Yeah. The key takeaway, I think, is that people could and should be much more excited about the positive implications of superintelligence, especially for science, physics, chemistry, biology, but also economics and astrophysics, and all kinds of other engineering tasks. I think there is so much more that can be done with better technology. And right now, I feel like a lot of people need, like, better marketing, not just for the future in general, but also, better marketing for technology and in particular for AI. And this book, should show even the AI skeptics, how much positive upside there is for AI, especially when it comes to inventing, new scientific discoveries. Swyx [00:02:09]: I think you quoted the techno-optimist manifesto from, Marc Andreessen, which I think was, like, beautiful in its, ambition and clarity and simplicity almost as well. Richard Socher [00:02:18]: I agree. Yeah. Yeah, you can disagree with him on some things, but, like, I think he’s right on the techno-optimism. Swyx [00:02:23]: Where do you think optimists get in trouble? Richard Socher [00:02:26]: Like, you shouldn’t have blind optimism. You should be very clear-eyed, like, especially when with such an omni, like, use type of technology as AI is, you need to think about the potential downside scenarios, especially when people use it for things that you don’t want them to use it for. It’s a little bit like the internet, and I feel like people are trying to regulate AI sometimes because of those potential downsides the way you would regulate the internet, if you were to say, “Well, because there’s bad content on the internet, like torture porn or whatever, like, we should just make it slower. That way, you can’t share the illegal content as quickly, or we should make the hard drive smaller so you can’t store as much illegal content.” But I’m like, “That’s not how you regulate that.” that’s like saying like we should regulate intelligence in the abstract. What you should regulate to avoid those downside scenarios, even as an optimist, are the specific applications. Sure, I don’t want, like, some AI surgeon to, like, practice some RL moves in my brain. It should be fully FDA certified. Sure, I don’t want any random startup to, like, drive on the highway, and cause a major accident. It should, like, have proper certifications before it’s let loose on the highway. But I feel like those downside scenarios, that some optimists sometimes maybe don’t consider enough are fairly easily regulated, compared to, what the doomers are worried about. Swyx [00:03:54]: It — Slow takeoff is part of the strategy as well? Richard Socher [00:03:57]: I do think, as excited as I am about, AI and its impact for society and, culture even, and certainly technology and economics and wealth and, health and all of those things, as excited as I am about all that, I do think the most bullish people on the AI hard takeoff scenarios overestimate how quickly things can move. There are hardware constraints. There are physical constraints about, the compute substrate. How quickly can you get enough, GPUs on? There are also constraints in the economy where there are a lot of industries that don’t require an insane amount of complex intelligence and complex capabilities. Like, if you think about jobs in, brands and, like, clothing and apparel and, like, handbags and stuff, superintelligence isn’t gonna make your fancy $10,000 handbag any fancier? Richard Socher [00:04:57]: It’s like that’s — It will have no effect on the economy. You think about travel and tourism. People wanting to see the pyramids, in Egypt, it’s not gonna change that much with AI. Sure, you can, like, generative a fake, photo of you and next to the pyramids. Swyx [00:05:12]: I can use Genie and, tour the pyramids in Genie. Richard Socher [00:05:15]: Yeah, exactly. But, and there’s so many industries, like logging and oil. You’re not gonna magically get 1,000x more oil because, like, sure, there will be robotics, like drilling and things like that could be done, but it’s not gonna 1,000x that industry in a, like, crazy hard takeoff scenario, both on the economy, and I can go on and on about all the other examples, where that, like food and so on, where that doesn’t necessarily change that much. And then, yeah, there are real physical constraints. And then there are, of course, like, people like, off-ramping from progress. That’s one of my concerns often is that I see people in, like, Europe and other, whole regions almost feeling like they. Like many people there wanna off-ramp from progress, period. And that will also slow down, like, more improvements. Swyx [00:05:59]: Yeah. We have this pulled up where, this is one of those things that, is very topical right now because now all the Frontier Labs are calling for the option to pace AI. They don’t say pause, they say pace. I don’t know if there’s there’s any take from you about, like, whether or not this will be effective. Pacing AI, Regulation, and Safety Incidents Richard Socher [00:06:17]: I think the downsides of trying to truly regulate with the full power of law what people do on their GPUs, would be worse than any of the concerns that they have. Like, it would be an crazy totalitarian state Richard Socher [00:06:37]: If every one of your GPU computes was known to some big government or multi-government agency. Richard Socher [00:06:44]: It’s like, it’s literally if you try to regulate intelligence, it’s trying to regulate thought, and that’s ridiculous, and it’s crazy. I think it is make — it is sensible to regulate some of the applications of this technology. Swyx [00:06:55]: Yeah. We had a bill, actual bill to regulate the number of flops in a model, and I’m like, “Okay, well-” Richard Socher [00:07:00]: Europe done it. Like, these guys have been successful enough with their fearmongering that all of Europe has regulated itself so much before it even had a proper AI takeoff because they listened to some experts who say, “We might all die if this technology has more than this number of flops.” And they’re like, “Well, we’re good. We wanna want people to thrive. Let’s not have technology that could have a small chance of all of us dying.” And so they regulated exactly those kinds of things in the EU. And so it’s, it’s very unfortunate that there are real implications for some people when others saying, “Let’s pace while they’re sprinting as fast as possibly,” “as fast as humanly possible towards that frontier themselves.” Swyx [00:07:43]: Yeah. It’s also not a global pause, right? Like, other nations are still accelerating at the same pace. Richard Socher [00:07:50]: Oh, yeah. Richard Socher [00:07:50]: You’d need a totalitarian world regime if you tried to regulate intelligence and GPUs and what people do on them. Swyx [00:07:56]: Any takes on the safety angles of this? So there was a drawback of Fable, a pause on 5.6 before it could be released. Recently, there was Hugging Face with the OpenAI cyber incident. Any takes there? Richard Socher [00:08:11]: 100 percent. I think these are serious issues of reward hacking, and clear failures, of doing proper red teaming or rainbow teaming. I don’t know if you saw this paper from Tim Rocktäschel and a few others, where one AI, is tasked to try to hack another AI and then they can go back and forth in an open-ended fashion to inoculate themselves from those. Yeah, this is the paper. It’s a really clever idea. Open-endedness, and evolutionary inspirations are, big for us at Recursive as well. And so I wish they had used more of that. And it’s clear that, for instance, the constitutional AI. I don’t know if you remember anthropic.com/constitution. You can pull it up and search for cyber right there. It says, “Hard constraint. Claude will never ever do cyberattacks, and that is a hard constraint in our constitution.” So here are the current hard constraints on Claude’s behavior. Richard Socher [00:09:16]: Number 3, create cyber weapons or malicious code that could cause human damage. Richard Socher [00:09:21]: And clearly, this whole constitution was fake. Like, it clearly isn’t being adhered to at all. Swyx [00:09:26]: Because Anthropic also found that they had in their testing Richard Socher [00:09:30]: They’re also. Like, they’re like, “Oh, well, other people are hacking now.” There are a couple things. One, you can make a sandbox very simple, and then it’s very easy to hack yourself out of a sandbox, right? But what I think it shows is that we’re currently in this state of AI where the reward engineer still has to do a lot more careful work, and where the AI, in most cases, is not very good yet at understanding what is meant versus what is being said. And so concretely, I think this will happen if we were to have this intelligence more easily accessible in a lot of companies. Imagine you run a service center and someone says, “Oh, here’s my CSAT score and my dashboard. Make this number go up.” It’s like, “Our CSAT score is so poor.” The intelligent AI will just be like, “Oh, sure. Like, I’ll just create 1,000,000 bots that call our service center and give a 5 out of 5 rating at the end, and the number went up just like you asked for.” And you’re like, “That’s not what I meant.” “I meant with our real customers.” The AI goes off and says, “Well, easy. I’ll just give a 1000 dollar gift certificate for every failed, whatever DoorDash Richard Socher [00:10:35]: Offer.” It’s like, “That’s not what I meant.” It’s like, “Well, but that is what you said.” And like, so I think clearly articulating what the rewards are is something we haven’t gotten very good at as humanity. And then clearly, the AI in these cases has not gotten good enough at understanding what we mean when we ask it and give it certain rewards. Now, what gives me hope is there are the first inklings, of this being better. I’ll give you an example like WhisperFlow. Full disclosure, I invested, in their seed round, but at AIX Ventures, but, WhisperFlow has gotten much better at writing what you mean and not what you say. And I think that is a sign of things to come. I think there will be more and more AIs as we make it more and more intelligent that will be better at being aligned with what is meant. Swyx [00:11:21]: Will it be done through a constitution or RLHF or Reward Hacking, Alignment, and What We Really Mean Richard Socher [00:11:23]: Clearly, constitutions don’t matter at all. Richard Socher [00:11:25]: It doesn’t work. And that was, I think, mostly marketing. I think we need to find better solutions for it. And I think at Recursive, we have a few very good ideas and some already Richard Socher [00:11:34]: Like, ways where I think we have a better grasp on it. I don’t think we’ve fully, figured it out yet, but, we’re thinking a lot about safety, and the more intelligent the AI gets, the more you want it to be aligned, the less you want it to think about reward hacks and try to do the right thing. Swyx [00:11:49]: I don’t know if we’ll touch on this topic, but I’m just gonna throw this question in here because it’s something that’s weighing on me. Alignment, let’s call it, is alignment to general humanity’s preferences, the median preference. Personalization is pinpointing what you want, and sometimes alignment can conflict because what you want is not what the general median population wants. How do you choose? Alignment, Personalization, and Cultural Values Richard Socher [00:12:12]: It’s a great question. Richard Socher [00:12:13]: I think you ultimately have to, of course, be aligned with laws. Like wherever your AI is deployed and needs to align with the law. I do think what AI often does is put this mirror in front of us and say, like, “This is what you’re looking like. Now I can amplify that a 1000 times. Is it still what you want?” and the truth is that different cultures made different choices. Like, in Eastern cultures, the greater good is often valued more, than the individual. Western civilization, we care more about individual freedoms and rights and the pursuit of happiness and so on, than others. And even there are gradations. There’s regulation versus litigation trade-offs. In the US, you first can often, not every time, like, FDA and so on does regulate some areas, but in many cases, the bad things happen, someone sues someone else, and then there’s a law based on that. In Europe, they try to often avoid any harm to anyone and regulate before. And both are, trying to do the best thing, but, some is more amenable to innovation than others. And so yes, you’re right. Like, I think ultimately each individual, each country, and humanity as a whole has to think about those values more, and then try to put them into laws. And that those are ultimately the constraints. And hopefully, different, societies, just like now with their AIs, will align their AIs to a different one so we have not just a monoculture of alignment. Vibhu [00:13:46]: Here’s a follow-up on this that I wasn’t expecting to ask. Do you have takes on open source, open weight versus who owns the intelligence? So, clearly not the biggest, fan of the constitution Richard Socher [00:13:58]: You had to do this in the topic side off. Vibhu [00:14:00]: But it’s fine. Vibhu [00:14:02]: Point being, any thoughts on who should own weight? Should it be open? Anything there? Open Source, Soft Power, and Who Owns Intelligence Richard Socher [00:14:06]: 100 percent. I am a big fan of open source. We’re gonna sign some various open source letters at, Recursive also. I think, even in the worst case attack scenarios, it is better to have more good actors have more different types of AI, accessible. I think, open source is a little bit a soft power type of thing, too. So I do think it’s good for the Western world Richard Socher [00:14:31]: To have an answer to that, out of China. I do think, when you watch a Hollywood movie, there’s — it’s like, I don’t wanna misc, diss all of movies, but there’s a certain sense of propaganda, right? You watch one side of things, right? Vibhu [00:14:46]: Oh, yeah. Have you seen Top Gun? Like, come on. Vibhu [00:14:48]: Like, it’s like half of it’s paid for by the US Army or something. Richard Socher [00:14:51]: Yeah. And so. And, I think that’s just natural. Like, but what’s interesting here is I think LLMs are essentially a similar type of soft power to movies and beyond, because they’re also, highly important for cybersecurity and so on. But one of their many aspects is that soft power of storytelling. Like, if, like a child asks an LM, like, “Tell me an inspiring story of what I should do when I grow up,” right? It’s like those are all these, like, subtle things. So I think it’s important, for Western world. I do love, individualism. I do think, despite, some of its flaws, like capitalism is the best way we have governed, found ourselves to govern, and so on. And so I do think there are various aspects that would be good, to have a Western open source answer, for LLMs. And, with Recursive, I can’t make the announcement quite yet, but we’ll Richard Socher [00:15:43]: We’ll be relevant in that space very soon. Vibhu [00:15:46]: Okay. All right. Exciting. I wanna bring us to Recursive. So outside of our tangents, you have a pretty deep background in the NLP space. You worked on, like, early embeddings, GloVe with Chris Manning, who was a previous guest on the podcast, You.com. What’s the history? How did you decide to start another company? From You.com to Recursive Richard Socher [00:16:06]: Yeah. So I’ve been excited about AI for over 2 decades now. I sometimes feel like it’s ancient history now. It’s BC, the before ChatGPT era. No one cares about all the religions that happened, before, Jesus Christ, and no one cares about the models that happened before, transformers and ChatGPT and stuff. But, like, it’s something that I’ve been deeply passionate about. I think AI is one of the most interesting things one could work on, period. I think language is the most interesting manifestation of human intelligence, too. And, at You.com, we eventually off-ramped from pushing, like the frontier of AI forward to mostly giving people, like, good search engines, search, APIs and answers over the web. I think that’s an extremely important part of intelligence, just knowledge and access, especially even, we’ll get there maybe later, if you wanna invent a eureka machine that invents everything for us, it needs to know how not to reinvent the wheel, proverbially speaking. And to know what has been invented, you gotta have internet access. So it’s the number one used, most used tool, in LLMs, agents, chatbots, and so on is web search. So I’m really excited for You.com to own that and grow really well in that with really large customers and so on. But it’s also not building frontier models anymore. And so I initially tried to do this within You.com and raise another round and so on, but you just can’t. You have to do a certain thing, and until you print enough money that you’re allowed to start a second thing within that company is really hard. At the same time, I had all these ideas. I put them into a book. I finished the book last year, and I was like, “It’d be really fun to work, on this myself.” I felt like with word vectors, and then prompt engineering and, ImageNet and larger language models for protein generation, not folding and so on, I, me and my teams have pushed the field truly forward. And I feel like we can do it again, here at Recursive. And in many ways, what I observed over the last, 20 years in AI is that whenever we replace some human part of the process of creating AI with a learned system, improvements follow. And so. We’ve done that taking out manual feature engineering, like in sentiment analysis. I don’t know if you remember these old days where, like there are linguists, and they’re like, “Here’s how you negate, and there’s a, like, regular expression.” Swyx [00:18:21]: I went to Penn where we — they had, like the WordNet Richard Socher [00:18:24]: That’s right, WordNet, all of that stuff. Yeah Swyx [00:18:26]: Original. They use, our grad students to label Wall Street Journal articles and, like, really construct a knowledge graph of Richard Socher [00:18:32]: There you go. Richard Socher [00:18:33]: And WordNet started, was part of how we started ImageNet. But anyway, so, like, it was really, like, fun, to do. But when we replaced all of that manual feature engineering with vectors and neural nets and just backprop through everything, it started to work really well at scale. And so then everyone started to do architecture engineering, and I was like, “ that clearly can’t be it.” Swyx [00:18:53]: You mean, neural architecture search? Richard Socher [00:18:55]: Like, manually, they would say like, “Oh, I’m, I’m doing sentiment analysis, so I have a special neural net that’s really good at sentiment analysis.” And then the machine translation community had a special neural net for machine translation. Swyx [00:19:06]: I see. Richard Socher [00:19:07]: The summarization people had their own stuff. And I was like, “That clearly can’t be it. We should unify all of that.” So I had 2 papers. One is called Ask Me Anything, and the other one was called DecaNLP. And DecaNLP eventually got cited, like, 5 times by the first GPT paper. And, to me, that was, like a really a big step forward. And then, of course, you had to combine this idea of prompt engineering with transformers and with language models, and you put it all together, you scale it up, which is also a huge amount of work. And then, the field progressed a lot. I feel like the next step and maybe the last step of that history and the arguably, success has a lot of parents, only failure is an orphan, like my version of that AI history, I do feel like in that history, you can think about, “Well, what’s the next way to automate?” And that is the AI research itself, like the human, process of ideating, implementing, and validating ideas. Automating AI Research and Recursive Self-Improvement Richard Socher [00:20:01]: And in our case, ideas for AI. Richard Socher [00:20:03]: And when you have AI then help you with that, it, by almost definition, becomes a self-improving AI ‘cause it now does research on itself. And there are lots of different misnomers. Some people think auto research is already recursive self-improvement. It’s Swyx [00:20:17]: Yeah, and you explained that in the talk Richard Socher [00:20:19]: Completely different. Richard Socher [00:20:19]: But, to me, it’s the most interesting thing that I could be doing, and I’m really excited with the co-founding team. What’s interesting is we have 8 co-founders in total, including myself. And so The Recursive Founding Team and Darwin Gödel Machine Swyx [00:20:31]: They are gonna bring it up. Richard Socher [00:20:31]: Nice. Yeah. And they’re all. I could talk about all of them if you want. Swyx [00:20:34]: Super stacked. Richard Socher [00:20:35]: Yeah. Just an incredibly talented group of people. And we all came to the same conclusion, but from very different directions. Like Josh Tobin, is our CTO. He ran, a bunch of different, projects at OpenAI, like, Codex and deep, research, agents and ChatGPT agents and so on. But before that, he also worked in robotics, and he saw the smaller simulations, and how it’s gonna be really hard to scale that in full generality. And so that’s, that was his angle coming to recursive self-improvement. We have Jeff Clune who’s been working in, like, open-endedness for a long time, together with Tim Rocktäschel. Tim Rocktäschel also built Genie 1, 2, and 3, which is, like the most exciting and most sophisticated, I think, still world model, anywhere. And so they both came from this, open-endedness angle. Jeff also, I think, published one of the most exciting papers in recent years about recursive self-improvement called the Darwin Gödel Machine. Super interesting paper. If we could, maybe pull it up really quick Richard Socher [00:21:35]: It would be, like, super interesting to see ‘cause you see Swyx [00:21:38]: By the way, I love how many paper citations. Swyx [00:21:40]: You’re, you’re giving people a lot of homework, which I like. Richard Socher [00:21:42]: Love it. Yeah. And so, like Caiming Xiong, a rockstar, we worked together at MetaMind and Salesforce Research together. Alexey Dosovitskiy invented the Vision Transformer, one of the most cited, papers in computer vision. Tim Shi is, like also a unicorn founder. Yuandong Tian led RL at Meta. So just like, yeah, really fun to work with them, and the next level of people are just incredibly strong, too. So it’s been a really fun ride so far. So the first figure, you see exactly these kinds of ideas, that, I think, yeah, inspired a lot of us and now more and more people, where you have this archive of different coding agents. They learn how to self-modify, evaluate, and then create these phylogenetic trees, of, yeah, different ideas. Swyx [00:22:28]: That’s one foundation. So that Darwin Gödel is an influence. Swyx [00:22:32]: Open-endedness is an influence. Any other trains of thought that feeds into Recursive that I’m missing? Influences: Open-Endedness and Learned Systems Richard Socher [00:22:38]: Going to replace manual parts of the process of building AI Swyx [00:22:42]: I Richard Socher [00:22:42]: More and more Richard Socher [00:22:43]: With learned systems. Yeah. Swyx [00:22:45]: Which, and, like, merging different fields into one general, architecture. Richard Socher [00:22:51]: That’s right. Swyx [00:22:51]: Okay. It seems like language models are already pretty generalist, right? Swyx [00:22:55]: Your next token predicting your reasoning. Was there a time that you thought, “Okay, these are good enough to have recursive self-improving machines”? Are Current LLMs Enough? Richard Socher [00:23:05]: It was clear to me that they will happen, within, like a year or two, and then it did exactly happen, like, earlier this year, right? Earlier this year, AI really went from not just being code, but being able to code. And that is a big unlock. It’s definitely making everything a lot easier than it was, before the beginning of this year. Swyx [00:23:24]: One question that I think a lot of people have is the current LLM paradigm enough? Or, like, let’s call it autoregressive transformer, with reasoning, whatever. Don’t you need something else, some big unlock, whether it’s world models, which Chris Manning is working on, or memory, continual learning, all that stuff? Or is it all of the kinds, and you think the current, let’s call it transformer architecture, is here to stay and that’s it? Richard Socher [00:23:48]: A lot of thoughts. So number one, I do think it would be great to have less of a monoculture in AI research. Richard Socher [00:23:55]: Like, if you look at, AI conferences now, I still remember the days in, like, 2010 when I tried to get my first neural net papers and NLP conferences accepted, and they just desk rejected them because, like, neural nets were something, quote, unquote, “We don’t do in NLP conferences,” and just, like, desk rejected. And it was very brutal in the first years of my PhD. Now I feel like it’s almost like the field switched to the other side. Like Richard Socher [00:24:17]: Someone should try some other weird, crazy ideas now that aren’t. Swyx [00:24:20]: There’s also a few. I really respect, like, people still working on, like, GNNs and, like tabular stuff and. Richard Socher [00:24:25]: Yeah. Like, someone should still, like, do novel out there ideas. At the same time, I think whenever people say, “Oh, LLLMs are. Like, this is the end for LLLMs,” they just don’t, like. LLLMs are also not the LLLMs of, like the past, right? Like, they are so much more sophisticated now. There’s so many more clever things that people are doing. It — There’s, like, different stages of training. You have the whole RL training, and you can take actions and, like all of these things where that can go really far. And then the folks that come from the neurosymbolic, direction say, “Oh, this will never work because they can’t do neurosymbolic reasoning.” It’s like, I think they’re underestimating still the ability for these models to code, and code is neurosymbolic reasoning, and these models can code incredibly well. And so I do think there are, of course, more and more ideas that will be needed and we’ll continue to have. We’re seeing, like, more and more interesting high-level ideas coming out of the AI itself, too. And with really deeply integrating the fact that these models are code and can code, that line — I don’t wanna give it all away, but, like, I think that line has a lot more to grow. But it’s still an LLM, right? Even if that LLM codes for you and then runs that code in some integrated fashion. World models, I’m personally less bullish on. I think if you run a robotics company, you’re gonna build your own world model. I think world models are super fun, and Tim Rocktäschel came to a similar conclusion after building the most interesting one with Genie 1, 2, and 3, which is gaming is a huge application for world models. Can see I sometimes got stuck in some games and, like, got a little overly competitive in the wrong direction. And so I understand games are fun, but personally, I’d rather work on science than gaming. And so, yeah, I think LLLMs, a lot more room to grow. Swyx [00:26:16]: Yeah. I think there’s some interpretation of world models that some people have where it’s like, well, it’s okay, yes, there is that gaming element. There’s this — there’s the embodied robotics element. But the other part also is just, the more abstract sense of LLLMs are just modeling output, but they’re not modeling the chain of thought, inside the human that has created the output. We can annotate it, of course, but, like, it’s, it’s always, like, this Plato’s cave reflection of a thing rather than the thing, right? Richard Socher [00:26:43]: It’s true. Richard Socher [00:26:44]: But I would argue that, and maybe we’ll get there in the 10, spaces of intelligence, but I would argue that even our projection, our eyes is a projection of the real world. And, like, we have only a very narrow, band of the electromagnetic frequency spectrum that we can observe with our puny little 2 eyes and so on. Swyx [00:27:01]: It’s good enough. Richard Socher [00:27:02]: It’s, it’s good enough for now, but, like the upper bounds of where it could be are so much higher. And, like, to map, the visual world the way humans see it is also not necessarily, like the end-all be-all for visual intelligence. And I would argue that language is still the most interesting manifestation of human intelligence. And while our visual cortex is certainly less sophisticated, than that of, certain animals all the way down to the mantis shrimp who can, have, like, 2 independent eyes, 3 bands, trinocular vision and each eye can see all the way to, like, floating temperatures in 4D and stuff. Richard Socher [00:27:36]: Like, mantis shrimp, you should look it up. It’s like Swyx [00:27:37]: Way OP. Richard Socher [00:27:38]: Super crazy. Swyx [00:27:39]: Yeah. ZeFrank, mantis shrimp. Swyx [00:27:41]: It’s the best video in the world on Richard Socher [00:27:42]: I love ZeFrank, yeah. Richard Socher [00:27:44]: Big shout-out to him. But, like, I think there’s a lot more room to grow, but none of these, other animals have language that’s as sophisticated as ours, certainly not in writing. And once you can write, you can, start thinking about longer term civilizations. All of that is language. Programming is much closer to language. And I would argue, and this is, like an important thing in the spaces definition of intelligence also, is that all of these spaces are highly correlated, but visual intelligence is neither necessary nor sufficient for overall intelligence. You can be blind and still be an intelligent human being. And an AI can be blind and still be quite intelligent too. Swyx [00:28:25]: We were gonna bring this Richard Socher [00:28:25]: Which doesn’t mean that you’re not more intelligent when you have it. Yeah. Swyx [00:28:28]: We’re gonna bring this up. I might as well — Like, we have a classification of 10 types of intelligence that you had at the end of your talk. So I’m just gonna flash this up now for people to cover this. I don’t know if, maybe we’ll put this towards the end. We’ll come back to this. I just wanna mention that, you do have a philosophy that I like when people do lists because then I can just go through this and then it gets — it’s educational for people. But let’s go back. I don’t wanna get distracted. But, so effectively, I’ll, I’ll, reinterpret what you said as Yann LeCun is wrong. And then we’ll just Richard Socher [00:28:56]: Don’t quote me as that. I’m, I’m good friends with Yann. I think very highly of him in many directions. Swyx [00:29:01]: But he’s wrong. Swyx [00:29:03]: You mentioned GPT-1, and I cannot let any, Alec Radford, mention escape. Did you talk with him when he was training GPT-1? Like, any historical, fun stories there that you might come up? DecaNLP, GPT History, and Scientific Gatekeeping Richard Socher [00:29:18]: I did not, like, meet him a bunch of times. I think we met maybe once or twice at some conferences. But, like, he has told, I think Brian, the first author of the DecaNLP paper, that it did inspire him, and he cited it five times in the GPT-2 paper. So, and that’s, like Swyx [00:29:36]: Yeah, good enough. Richard Socher [00:29:36]: Very clearly said, like, this was the first instantiation where they showed in the DecaNLP paper, McCann et al, that you can just phrase every single NLP problem as here’s some prompt, text context, here’s a question and task description and here is some output. If you just do that enough, you can have one unified neural network model, which, by the way, also had all kinds of interesting attention mechanisms. There are slightly different formulations to the transformer. I think came out the same year, plus/minus a few months. And then you can unify all of natural language processing into one neural net. That is the core idea. Swyx [00:30:14]: And this was as opposed to at the time, LSTMs and what have you. Richard Socher [00:30:17]: LSTMs, but also, like, people being very stuck in thinking about one model per task. In fact Richard Socher [00:30:25]: It’s, it’s kinda crazy, but the DecaNLP paper was publicly reviewed as, like, open, OpenReview. It was an ICLR submission. And, in it, you will see, how the whole community at the time thought about this. So, like Swyx [00:30:43]: Some great contributions, but more work needed. Richard Socher [00:30:46]: So look at, like, search for not even for humans. Just scroll it up here. Like, question answering is not a unified phenomenon. There is no such thing as general question answering, not even for humans. And this is like, really, you replace your brain with a different brain a different neural net when you answer, like, different kinds of questions. It was unfathomable to the experts at the time that you can have one unified neural network that would answer all of these different questions. They are saying, “No, all of these questions require very different systems to answer, and trying to pretend they are the same doesn’t help anyone solve any problems.” That’s what it says right there, right? That’s how hard it was to fathom. And now, of course, people, when I say, “Oh, we’re gonna invent prompts,” people are like, “You can’t even invent prompts.” It’s such an obvious idea to have one neural network that, of course, does everything in NLP. Richard Socher [00:31:37]: But at the time, it was, like, extremely controversial, and the paper got rejected. And the sad thing is that it got rejected so hard and they were so certain that we stopped going on our list of things to try. And the number 2 or 3 on the list of extensions for this paper was add language modeling as another task. And then we could have, and that would have accelerated the timelines, in 2018, like, even further for humanity. But we got so crushed, and we were like, “Okay, maybe we’ll just work on some of our other ideas for now and, like, come back to this later.” Yeah. Swyx [00:32:09]: How can we design a review system that rewards non-consensus? Richard Socher [00:32:14]: Honestly, I started to feel like arXiv is such a gift to humanity. With arXiv, you should just put your paper out there. Swyx [00:32:24]: Is it pre-preprints? Richard Socher [00:32:25]: Let — And honestly, I think Twitter X, people like you who pick up interesting papers, that is a better filter than the experts. Let everyone, like, have access. Now, of course, there are some downsides, which is, like, if you’re super unfamous, you have no Twitter following Richard Socher [00:32:41]: You don’t wanna be on social media or whatever, you write a good paper, maybe someone, somehow no one notices it. But I would argue that if you just tell, like, 10 of your friends in your community about a paper and it is a really significant breakthrough, someone is bound to talk about it again. And, so I think science needs less gatekeeping. And, even though ICLR, with Yann LeCun, who started it, as one of the co-founders of ICLR back in the day, he also wanted less gatekeeping ‘cause he too was rejected for many years together with Yoshua Bengio and Geoff Hinton with all their early deep learning and neural net papers ‘cause it was just not the hot thing. And so ICLR started with that, but then it also started gatekeeping a little bit themselves on various ideas. So I think less gatekeeping, more open, and then allowing people to say, “Look, even if this is just on, or, quote, unquote, ‘just an archive,’ if it has like 1000 citations, it’s a legitimate paper. Doesn’t really matter where you published it.” Swyx [00:33:34]: And I agree with that. I do think it’s sad that I’ve heard that grad students have to do, like, how to Twitter, seminars to each other Swyx [00:33:43]: Just because it’s so important for publishing these days. This person is just reflecting the sentiment at the time. Richard Socher [00:33:49]: That’s right. Swyx [00:33:49]: But it’s Richard Socher [00:33:50]: I think it’s Swyx [00:33:50]: It affected you so much Swyx [00:33:52]: That you stopped work on it. Vibhu [00:33:53]: The sentiment also came out of some of the research, right? Like, the original BERT paper was trained, and towards the end of the paper, they’re like, “Okay, throw off the last head, train specific iterations for Vibhu [00:34:05]: Extractive summarization add a head for this.” Like, you should do task-specific stuff. These are, like the authors that wrote Attention, wrote BERT, telling you this is what you’re meant to do. And, like the training tasks were also very odd. They’re like Vibhu [00:34:16]: The — “We know that the model overfits to this weird mass language modeling. Throw away this part and just do specific models,”? Richard Socher [00:34:23]: Exactly. And, like, we had to try — come up with all clever ways of, like attention and pointers and so on to get the neural network to be able to do all of these tasks. And then some of them were better than state-of-the-art, some weren’t, but we were like, “But it’s still in one model.” I thought it was really cool. Really interesting. Swyx [00:34:38]: I was gonna move on next to Tim and open-endedness. He was head of open-endedness at Google. Open-Endedness, Rainbow Teaming, and Self-Set Goals Richard Socher [00:34:42]: That’s right. Swyx [00:34:43]: I don’t know what that means. Swyx [00:34:44]: But he did a lot of talks. Richard Socher [00:34:45]: Genie 3 is one of the ways that Richard Socher [00:34:47]: Rainbow teaming, yeah. Swyx [00:34:49]: So I first saw him at — speaking of ICLR, I first saw him at ICLR when he talked about open-endedness. He’s he’s done a few talks. Can we define what is open-endedness for people who have never been exposed to the problem? They are like, “What do you mean? I thought the only goal of AI is to optimize against a benchmark or.” Richard Socher [00:35:04]: That’s right, yeah. It’s a, it’s a fuzzy term because there’s so many different instantiations of open-ended, thinking. But, one way I often describe it, and certainly, Tim and Geoff Hinton would be even better at describing this, but it’s a suite of methods that is more inspired by evolution than, very specific rewards. So in that sense, it thinks more about environments, about co-adaptation. And so a concrete example is in the cybersecurity and LM safety space where you have one LM that tries to attack another LM to say something unsafe. Swyx [00:35:40]: Yeah, the rainbow, yeah. Richard Socher [00:35:40]: And now the environment is the 2 having a conversation and now they co-adapting, right? They’re like one makes a better attack than the first one inoculates itself somehow, like uses that as training data, makes it so it’s harder to say something unsafe based on that. And then as the attack stops working, the attacker now tries a different angle, right? Richard Socher [00:36:00]: And that’s why it’s not just red teaming, but they’re called rainbow teaming. Swyx [00:36:02]: So, like, don’t tell me how to do things. Let me just figure it out myself. Richard Socher [00:36:05]: That’s right. Think about the environments that you wanna use. Think about the rewards at a high level that you wanna, inspire towards, and then let the AI try out many more ideas in this interplay between sometimes humans, but also sometimes other AI agents. Swyx [00:36:22]: Yeah. I worked open-endedness into a model that I have been working on. It was the keynote for AI Engineer where you start. You, we have the token loop, we have the agent turns, and then we have goal. And I feel like the way that you’re describing open-endedness is still somewhat of a goal. Like, please attack this, Swyx [00:36:41]: Other agent. But, to me Richard Socher [00:36:42]: Yeah, you set the rewards. You set the environments. Swyx [00:36:44]: The loop that makes the other loops is. What if the agent can set its own goals? Swyx [00:36:49]: And is it, is that open-endedness? Like, you don’t give it a goal. Just, like, be a sentient being. And maybe sentient is a very loaded word Swyx [00:36:57]: But just set your own directions. What do you think you should do? Metacognition, Subjective Goals, and Measuring Intelligence Richard Socher [00:37:01]: I love this direction. I think this is one of the 10 spaces of intelligence, that I clump under metacognition and thinking about thought. Richard Socher [00:37:08]: And it’s an interesting one. Whenever people say, “Oh, AI is like, this is, it’s gonna stop from here. It’s not gonna get that much better,” and blah, I’m like there’s so many different spaces of intelligence that we haven’t even started exploring yet and hence have made very little progress on. And there is an interesting, connection to economics and, capitalism. Like, it doesn’t make sense for a company to build and spend billions of dollars building a model that instead of following the rewards and objective functions you gave it, may come up with its own objective functions and its own goals. Richard Socher [00:37:46]: Right? And then imagine you’re like, “Okay, I spent billions of dollars. Now go develop this new battery, material for me and answer all my emails.” And it’s like, “Nah, I think it’d be more interesting to evaluate the molecular composition of the atmosphere, on Jupiter.” Richard Socher [00:37:59]: And you’re like, “That’s not what I paid you billions of dollars for.” And so no one’s working on that for good reasons. And then also, understandably Swyx [00:38:07]: It’s not useful. Richard Socher [00:38:07]: It’s not, it’s not useful, and it could get a little bit weird, right? What if the AI does start to really have thoughts on its own, and what if we don’t like those thoughts, right? And so it requires a whole different way of thinking about it. I had a great conversation with a good friend of mine, Sam Gershman, who’s a neuroscience professor at Harvard, and, like, we just jammed on this a little bit on, like, what are the best meta goals. And, I do think, like, knowledge-seeking is a really good one. I’m currently thinking also about, like the ultimate measure and unit of intelligence broadly construed, and I finally have some. It’s still too early to share it. It’s not. I haven’t fully baked the thoughts yet. Swyx [00:38:44]: Like some replacement for IQ. Richard Socher [00:38:46]: IQ is such a terrible definition, right? Swyx [00:38:48]: Elo. Richard Socher [00:38:48]: It makes no sense. Yeah, Elos are terrible, too, because it’s always just like me versus others. Richard Socher [00:38:53]: But, like, you can be intelligent and not constantly compare yourself to others? And so, yeah, there’s no, like. In fact, a lot of these definitions we have, which I briefly mention in my book, too, these definitions create sometimes explicit and sometimes a more implicit anthropic bounds. No dis to the company Anthropic, but just, like, this idea that your intelligence is like getting 100 out of 100 questions right on this IQ test. Well, if that’s your definition then you can only be at 100 out of 100. Where do you go from there, right? So you see a lot of these, benchmarks that people are working on they, increase, they get close to human, maybe sometimes Swyx [00:39:30]: It’s like an S-curve Richard Socher [00:39:30]: Slightly above human, and then it’s flat. Richard Socher [00:39:32]: It’s like, ‘cause that’s your. If your definition is only that so tied to humans, you’re only gonna get to just slightly better than that. So I think metacognition is a great example of that, where we’re not even yet allowing the AI to think. We’re not working on it very much, and hence there’s very little progress in that. Profit Maximization, Real-World Environments, and Reward Design Swyx [00:39:49]: Yeah. Well, we’ve interviewed Andon, which I think, has been working on the most open-ended, benchmarks, which is just real-world, money. Swyx [00:39:57]: Arguably, telling an AI to profit maximize is a bad idea. Swyx [00:40:03]: But they are doing it. Richard Socher [00:40:05]: I do think you don’t want that super. Like, you don’t want a superintelligence to have a ton of access to all kinds of tools and so on and then just give it that without some very careful reward engineering. ‘Cause it’s like, I just buy a bunch of defense stocks and I start a war. I make money. Like, it’s just like, it’s a tricky situation, right? You just buy a bunch of stuff, short basic goods for people, and you create some weird famine, like, issues. Like, yeah, there’s a lot of constraints you should put onto a trading system. Vibhu [00:40:35]: It’s a fun measure, though, ‘cause, the bounds are very capped to where we’re nowhere close to them. Like, in Andon Labs, the model’s like, “Oh, it’s Saturday, maybe I just close the store today.” “Someone’s off. It’s okay. We’ll just close the store.” Swyx [00:40:51]: It’s using Claude. Vibhu [00:40:52]: Yeah. But Richard Socher [00:40:53]: Yeah, no. I’m not, I’m not arguing against it. Just, like as you get more and more intelligence, you wanna be more and more careful with that as, like an open environment, ‘cause the environment then is all of Earth. Applying RSI to Science and Invention Swyx [00:41:02]: Yeah. Okay. For recursive, not strictly necessary, right? Because, like, if your goal is you make a machine that, like, invents the other things, then, like, just solve, the science things Richard Socher [00:41:12]: Knowledge discovery, yeah. Swyx [00:41:13]: Solve machine learning research and discovery and all these things. Good enough. Richard Socher [00:41:16]: And eventually, so, our goal, I haven’t really. I don’t talk about it that often because it is a few years out, but our goal is once you have a recursive self-improving superintelligence, you then want to apply it to the most important problems. And I think a lot of those are in science and technology and broadly construed inventions, and those inventions in, physics to create better, cheaper energy with fission or fusion, in chemistry and to create better materials and better batteries and, better solar cells and so on. In biology, there’s so much, like, I think soon to be low hang- lower and lower hanging fruit because of AI, because of protein and generation, not just folding, but generating new proteins like we did in ProGen many years ago. Like, so much positive impact we had if you take that superintelligence and you apply it to science. Swyx [00:42:04]: I do fundamentally believe that. There’s a lot of approaches, though. You’re not the only team trying and NeoLab trying. Swyx [00:42:09]: There’s, like a lot of. Especially the physical sciences as well. Richard Socher [00:42:12]: And that’s good. Yeah. I do think that physi- like the reason we are only doing it in a few years is that it’s a little too early right now. Robotics is not quite there yet. The AI is not quite there yet. But I’m fairly confident in 3 to 5 years, all those constraints will be gone, and then applying to real physical robotics experiments and so on, like true robotic process automation Richard Socher [00:42:33]: Not the traditional RPA sense, but, like, having robots run experiments for you will be totally there. Yeah, it’s gonna be great. Swyx [00:42:40]: Just to call back to something that you said early on about slow takeoff, you said that, like, while really the substrate that is limiting factor is, let’s call this chips, and semiconductors and all these things, and you have race funding for that and, you are investing a lot on that. But have you done the math on, like, is it even- Achievable and, like, what is the, industry concentration needed in order to achieve, like, scale? Compute, Slow Takeoff, and Changing the Bitter Lesson Slope Richard Socher [00:43:05]: Right now we know that, like, roughly, like a 1000 GPUs cost quite a lot of money. Richard Socher [00:43:11]: Right? If you wanted, like, 10s of thousands of GPUs, you’re, you’re talking billions and billions of dollars. If you say, like, one GB300 is, like, you could eventually create models that are, on that substrate, like are close and similar to human intelligence. And you want, like, thousands and thousands of, AIs to think about really hard problems, in a similar fashion to humanity. Like, yeah, that-that’s, that’s a lot of money. You do the math. It’s like a lot. We don’t have that amount of money right now anywhere to, like, build that. Now, things can get more efficient. You will have, I think, soon better algorithms that won’t be, and better hardware that won’t be as energy-hungry, and so on. Our human brain does quite a lot of flops with much less energy. Swyx [00:43:56]: 20 watts? Richard Socher [00:43:57]: That’s exactly right. Yeah, that’s the number often that’s quoted. And, like, I think more, inventions will happen there, that then will accelerate the takeoff even further. Swyx [00:44:08]: One thing I always try to reconcile when talking, like, with new lab founders is, like, you’re fighting Bitter Lesson all the time. You have to show initial progress, then you unlock the next tier of funding, then the next tier, then the next tier. Richard Socher [00:44:20]: Which unlocks larger model categories. Swyx [00:44:22]: Like, fundamentally, is that true? Like, are you fighting Bitter Lesson? Are you — will we have a way in which, like, no, we’re changing the slope in some fundamentally different way? Richard Socher [00:44:31]: I do think we are changing the slopes in fundamental ways by making AI much more efficient, both in terms of the training as well as the inference. Richard Socher [00:44:43]: Yeah. I think we will — When you allow AI to do the work that it takes other labs thousands of people and years to do, I think we’ll be able to get it down to weeks, and that will be much cheaper Richard Socher [00:44:53]: And hence, more affordable, accessible to others and so on. Swyx [00:44:57]: Yeah. You’ve shared initial results on that, Swyx [00:44:59]: Which, like, conveniently OpenAI has also done to their GPT-5.6, so we can talk about it now. Richard Socher [00:45:04]: Yeah. Yeah, so these are Swyx [00:45:06]: Let’s recap what you’ve done. Early Recursive Results: NanoChat, NanoGPT, and SOL-ExecBench Richard Socher [00:45:07]: Maybe, just a quick recap here. We built, this, system that isn’t the full, even the full RSI system in its glory, but it is a first baby version of this. And then, we don’t wanna just have it internally and not show anything and, just show some people of what’s possible. And so we applied this to these 3 different tasks. One is NanoChat, by my friend Andrej Karpathy, just, like, train a small language model to get, really low bits per byte. And, like, hundreds if not thousands of people, used both their agents and themselves to try, to get to that, and then they got to 0.937. We literally took our system and got to a much lower, bits per byte, much faster within, like, I think less than 2 days. So we took this thing, applied our system to it, and less than 2 days later, we have — we outperformed every human and their agents, in, have ever worked on this. Same with NanoGPT. And then we’re like, well, let’s, apply it to something that’s even more relevant, to real people and to the Nvidia ecosystem and applied it, to, SOL-ExecBench. And maybe you can scroll down to some of the, images. They’re, they’re kinda fun to see. But yeah, like, one you see has made some real inventions that weren’t just hyperparameter tuning. Like, inventing hash tables and so on is quite clever. We have even better results now. Swyx [00:46:34]: What do you mean inventing hash ta — You didn’t invent hash tables. Richard Socher [00:46:36]: Of course we didn’t invent, like, hash tables. In the grand scheme of, like a hash table, it’s like a super basic primitive in computer science. But to use it, for language modeling in this scenario inside a transformer and so on and to combine these ideas and put them together, that has then eventually also been invented, but there was a knowledge cutoff, and we did check that it didn’t have access to that externally. We talk about this a little bit. If you scroll to the next figures, this is also an interesting one in that when you start from a really basic, poor, like, vanilla transformer, then we still outperform all of the community together. But if you start from the human seed from an expert like Andrej, then you get even lower. So the human seeds from which you start do still matter. So that was an interesting insight, in my eyes, on this. And then as you go, like, how long does it take to get to these models, to get to similar performance? It’s much faster. And then a similar thing happens with the speed runs here where, people have worked on this for quite some time, and the model still was able to train a model more quickly. Why do we care about it? Well, speed of training is part of the equation of the cost, and ultimately, you wanna have the most intelligence per dollar, right? And so speed and quality are big parts of that. And, the, Swyx [00:48:00]: Yeah, the way I put it is, for people who don’t understand they look at the chart, they’re like, “Cool. What does it mean?” if you have, like a billion-dollar cluster and you can shave off 10%, that’s 100 million dollars. Richard Socher [00:48:12]: That’s exactly right. Swyx [00:48:13]: How much is that worth? Richard Socher [00:48:14]: Exactly. So when you click, when you look at, like the kernels, these kernels, yeah, for the non-experts, like these kernels are like, used in all the models. Every time you use an Nvidia GPU, you interface with that GPU through these kernels. And so here you see, the leaderboard best, and when it’s recursive, and it’s there are only a handful of kernels, in this whole benchmark where we weren’t the best. And so to me, this is, like, really exciting, ‘cause it makes. It just showcases what this can do. And again these weren’t like. We didn’t, like, spend months or years, like, developing. In fact, in particular for kernel, CUDA kernels, like, we don’t even have really deep. CUDA kernel experts in the team. And our system, that’s the beauty. The system just did all of these things. We didn’t invent this. And when we open source and release, things in the future and models in the future, like, it won’t. They won’t be the best in their, category or class or whatever because we’re so smart, but it’s because, we built a smart AI that does it for us. Reward Engineering and Good Auto Research Vibhu [00:49:14]: Do you have anything that you’ve learned from how to guide good auto research? A lot of it also builds on human background, right? It’s not just as simple as just, “Hey, go optimize this.” Vibhu [00:49:23]: But we do see it again and again, right? Like some of the Erdos problems, frontier math is being solved by people. And when they do a write-up, they’re like, “Oh, I’m not a mathematician. I have no background in this?” “I saw some tools and I made it work.” Swyx [00:49:35]: While you’re watching the World Cup, you’re like Swyx [00:49:37]: “This proves some conjectures that’s going on.” Vibhu [00:49:40]: Yep. Any learnings from Richard Socher [00:49:41]: Yeah, there’s a Korean conjecture was. Yeah, that’s pretty cool. Swyx [00:49:44]: To summarize, tips for good auto research Swyx [00:49:46]: Versus bad auto research. Vibhu [00:49:48]: How did you build the recursive? Richard Socher [00:49:49]: Yeah. So without giving away all the secret sauce, maybe some things that are probably obvious to the experts but might still be interesting to some, folks is, like, reward engineering is one of the most crucial bits, especially, in order to avoid reward hacking. So you have to be really clever about avoiding. ‘Cause as your AI gets better and better, it will get better and better, at finding weird like, special cases or counterexamples and things like that. And so I’ll give you an example. Like, when you ask to, like, make these 100, lines of code faster, and, how do you define fast? Well, you have one line at the beginning that says, “Start your stopwatch,” and one line at the end, “End the stopwatch,” and then, tell us how much time, progressed. And so, well, the simplest way is you just put that line that ends the stopwatch, right Vibhu [00:50:39]: At the start Richard Socher [00:50:40]: At the start. And then boom, it’s now faster, right? So this isn’t like this, like, super evil AI. It’s just, like a very simple, dumb reward hack. And so you have to just very carefully think about all the different angles there. And then I think the longer time horizon the tasks are the harder it gets and the more interesting and clever you have to be to still use these kinds of ideas for it. But yeah, I can’t give away too much there. Vibhu [00:51:05]: It seems like rubrics are taking a good spot in that, where for unverifiable domains, you have rubrics, you have a model breakdown, judge’s criteria along the way. Swyx [00:51:14]: Yeah, it’s a form of verification Swyx [00:51:16]: Once you got enough rubrics. Richard Socher [00:51:17]: Yeah, everything. I said this a long time ago. That’s why I’ve never been that impressed that AI can play games, ‘cause I’m like anything you can simulate and/or verify, you can have infinite training data for Richard Socher [00:51:29]: And hence, like, AI will solve it eventually. Swyx [00:51:32]: Looking for games where you can do auto domain distribution. So this is a game that nobody’s trained on ‘cause it’s a new game. Swyx [00:51:38]: And you can start gaming, you can start to play. So I’ve been building this and cloned this in person and it’s just been self-play. I’ve had about a billion positions evaluated. Games, Self-Play, and the AI Economist Swyx [00:51:48]: And, I wanted to do the AlphaGo thing of self-play until you get better, right? Swyx [00:51:53]: Like, which is like. This is not even LLM AI. This is just classical game AI. Swyx [00:51:58]: But, I think that the. And, but I set GPT-5.6 to auto research it because, like, I don’t wanna hand- handle any of this. I expect, the AlphaGo process to be, like, fully in the weights by now. Swyx [00:52:10]: It is not. It is. It, like, immediately leveled off very immediately until I human play tested it, and then I, like, called out obvious mistakes, and then they were like, “Oh, yeah. Okay.” And then it just dropped again. Richard Socher [00:52:22]: Yeah. Yeah. Yeah. Swyx [00:52:23]: And like, no amount of, like, think different, think more creatively, give me 8 different directions, any. No amount of prompting got it. Richard Socher [00:52:31]: Interesting. Swyx [00:52:31]: Like, you had to, like, RL against a human to Swyx [00:52:35]: Do it. So I, that was my. And by the way, Bean always wins if you. If anyone watches, Reese Ender’s Game. Vibhu [00:52:42]: And you put quite a bit of work into the guide for the AI. Like Swyx [00:52:46]: A lot Vibhu [00:52:46]: So the game you stack tiles. There’s some rules. You wanna capture the most area. You have, like a whole 50-pager on every rule. Vibhu [00:52:56]: You fed that in. It couldn’t, it couldn’t handle it that well. Richard Socher [00:52:58]: Yeah. It’s so funny that this reminds me of the claim territory and stuff of a paper we did in 2018 called The AI Economist. If you search for AI Economist Salesforce, we had a video we can play. It was an economic sim. Richard Socher [00:53:12]: So the idea is you have all these economic agents. They just wanna optimize their own utility function, which, is, collect resources that make money. And you can sell resources like wood, and then, over time, as you collect more, enough wood, you can build houses, you can trade with other agents, and you can use the houses then also to block off resources Richard Socher [00:53:35]: From other agents. Richard Socher [00:53:36]: So there’s, like Swyx [00:53:37]: Big strategy Richard Socher [00:53:37]: Competitive play and strategy Richard Socher [00:53:39]: And so on. And the point was that we wanted to understand what is the best way of taxation and subsid- subsidization to optimize an economy. And this research has not yet had its GPT moment, but I believe that countries like Singapore and others should and will eventually use this to, instead of doing, like, partisan politics and, like, special interest politics of, like, who donates the most to your campaign and stuff, you say, “Well, here, I wanna help the middle class,” or whatever you might say is your objective as a politician. And then people say, “Okay, well, how do you wanna do that?” And it’s like, “Well, here’s my fiscal policy. Here’s how I will change the taxes and pay these people,” and so on. And then you can put that into a simulation and you run that attempt from the politician against billions and billions of years of other strategies to try to achieve the goal that they set out to do. Richard Socher [00:54:36]: And then you can say, “Well, if that was your actual goal, then here is, billions of years of a strong simulation that would suggest that you try other ways of doing it, and maybe this the taxes and so on and this these tax brackets and so on.” And this is how you avoid gaming ‘cause these agents also try to reward hack to not pay their taxes and Richard Socher [00:54:55]: And so on. I thought this paper was super interesting. Unfortunately, similar to the first paper on, prompt engineering- The economists are like, “We don’t know any of this math.” It’s just like Swyx [00:55:08]: It’s not even, it’s not even math. It’s just we don’t trust your simulation. It’s not about math. Richard Socher [00:55:12]: It was — I, they just desk rejected the thing. And it’s like Richard Socher [00:55:15]: It’s like they didn’t even give us, like, clear like, clear signals. But, like the world of economics unfortunately doesn’t have proper Swyx [00:55:23]: Oh my God. Richard Socher [00:55:24]: Yeah, it doesn’t have proper, benchmarks. So you cannot be. Like, eventually, why did neural nets win? Not because people loved it. Like, they had all kinds of beautiful integrals and graphical models and stuff, but it just worked better. Richard Socher [00:55:36]: But in economics, it’s hard to prove Swyx [00:55:38]: So empiricism versus. Yeah. And I do have a bit of that econ background where, like there’s a lot of physics envy where you wanna write the general equation for an economy, versus just simulating it and using an evolutionary approach. Swyx [00:55:51]: Vibhu was thinking exactly what I’m thinking, is didn’t we have the GPT moment with small, Smallville? Richard Socher [00:55:56]: Yeah, I love this. Hello. Yeah, they Swyx [00:55:57]: As well, Dune, Joon just announced. I don’t know if you guys are involved. Simulations, Economics, and Policy Vibhu [00:56:00]: Simily there. Swyx [00:56:01]: Simily, that they’ve Richard Socher [00:56:02]: I wish we were involved. We’re not, yeah. Swyx [00:56:04]: Yeah. I had a couple simulation-based talks at AIE, so if people wanna look up what the state-of-the-art there, a lot of people are exploring this. It is Vibhu [00:56:13]: Proven out. Swyx [00:56:13]: Yeah. We also had a podcast with Mikhail Parakhin from Shopify, who is using simulation for commerce. Swyx [00:56:20]: Which, will simulate, like, your trajectory and, like, predict what changes, you make to your commerce journey will affect in your sales and all those things. Richard Socher [00:56:27]: I love this. Yeah. It’s really hard to simulate an entire economy, right? You have to make some simplifying assumptions. Swyx [00:56:32]: It’s just, everything’s, “Oh, LLLMs is very expensive.” Richard Socher [00:56:34]: Exactly. Swyx [00:56:34]: And I’m just like, “Am I gonna do this 8 billion times?” Like, come on. Richard Socher [00:56:37]: Exactly. Richard Socher [00:56:37]: But, I feel like countries like Singapore that really wanna just objectively do the right thing, have very technical leadership and so on, like they might like, eventually really try to simulate their economy. And you have to make some simplifying assumptions, but it gets really interesting ‘cause you can also say if your assumptions are such that all people would work hard if you let them, and they have the free. And then it turns out you have to make assumptions. Like, well, some people’s utility function of, like, how many hours in a day do they wanna work are different, right? And then you can start to disagree on the assumptions that go into the simulation. And then once you say, “All right, now we agreed on those,” or we have different views of what people are like at different, distributions and whatnot, then there are different outcomes, based on your goals. And then, of course, humans should choose what are the goals. In our case, it was productivity multiplied with equality, which, has some issues, but it’s, like, not totally unreasonable. Swyx [00:57:29]: Yeah. Just a comment on Singapore, ‘cause you probably have no idea, but, I am Singaporean and I’ve, been involved in the Singapore AI Council for making these things. The main reason they won’t is because they’re very conservative. Swyx [00:57:42]: And, I try to view it as the. There’s a founder-led country. When you start a country or you start a company and it’s founder-led, and you can do whatever you want because it’s your country. Swyx [00:57:52]: And then there’s manage- like, professional manage- managerial class, which is now. That’s, that’s what Singapore is. So they wanna. They always wanna see someone else do it first. Swyx [00:58:00]: And. But, like, everyone in the West views Singapore as like, “Oh, it’s a small country. You can do whatever the hell you want.” Like, Singapore doesn’t do that. Swyx [00:58:07]: So, like, someone else has to take the charge there. I’m just gonna do one question on the simulation thing, and then I don’t know, we can probably move on. Mode collapse, right? Like, LLLMs do not model the decision of humans. Spamming it out 8 billion times is not gonna help you model humanity. What do we do? Mode Collapse, Persona Simulations, and LM Arena Richard Socher [00:58:25]: I do think, you have to be clever about prompting each one individually. Richard Socher [00:58:31]: And I think that will help you get stuck into different modes. And in a weird way, people also get stuck in different modes? Like, there’s a lot of people, like, don’t teach an old dog new tricks thing. Like, once people are stuck in their ways, the older they get, the harder it is for them to think new ways. And there’s this, I think, comment, I forgot who said it, but it’s like, everything that was invented, before you were born is natural. Everything that is invented when you’re 20 is cool. And everything that’s invented after you’re 60 is, like, unnatural and an abomination and weird. Richard Socher [00:59:02]: I feel like that’s. It’s, it’s true for a lot of people. Like Swyx [00:59:05]: Yeah, it is a fashion and, I think people will do it. Tencent had a billion personas paper that gives a good data set for prompting, simulations if anyone’s looking into this, on the podcast. They just had, like, “You are a 30-year-old grocery store clerk. You are a 50-year-old professor.” Swyx [00:59:24]: And then just do a billion of those. Richard Socher [00:59:26]: Checks out. Yeah. Swyx [00:59:26]: So then you just use it. Richard Socher [00:59:27]: I’m, I’m shocked how well a lot of these things do map to ultimately similar statistics to real experiments. Yeah. Yeah. Vibhu [00:59:36]: I think it’s also good stuff for people to try that when they get into research, right? Like, we’ve seen train a model only on data before a certain date and see how well it extrapolates out. Do the same thing, right? So, see, do people code more with better coding agents? Can a model that hasn’t been trained on this figure that out without web access, right? Extrapolate out. Test these things. Richard Socher [00:59:56]: Just today, I think LM Arena published a interesting result where they were able to create a model now to predict your ranking. Swyx [01:00:03]: Wait, based on what input? Richard Socher [01:00:05]: Your model. I guess you give it your model, and it predicts the Elo score. Swyx [01:00:08]: I see. Okay. Sure. Richard Socher [01:00:09]: It’s surprising. Richard Socher [01:00:11]: Their whole raison d’être is like, oh, like, we help you compare these models. Yeah. Swyx [01:00:16]: Yeah. This team, they- they’ve done a lot of work, and they have the most data to do this, so why not? Richard Socher [01:00:20]: Right. Yeah. Richard Socher [01:00:21]: That’s probably right. Swyx [01:00:22]: When they were coming out of UC Berkeley, they not only had LM Arena, but they also introduced a routing project Swyx [01:00:27]: That would route based on LM Arena. Richard Socher [01:00:30]: Makes sense. Swyx [01:00:30]: And I don’t think that ever came to pass, and I’m curious why. I never got to ask them about it. Swyx [01:00:35]: ‘Cause, like, it’s. It was like, oh, yeah, clearly that’s your business model. You will become a router. Swyx [01:00:38]: And they never became a router company. AI for AI: Kernel Optimization and Inference Efficiency Swyx [01:00:40]: Weird. So that. I’ll just, put that out there. We’re gonna talk about GPT-5.6, self auto research thing if you have anything. I should also mention in your list of, kernel optimization and on the track that you spoke at, we also put Zhengyao Wei from Vico, who was also number one in the Parameter Golf Challenge, which is an OpenAI hiring, challenge. Swyx [01:01:05]: Which is also a very similar story. I think we’re gonna just see this all the time, where Swyx [01:01:09]: Humans optimize a thing a lot, and then some Richard Socher [01:01:12]: AI team comes in and just becomes number one. Swyx [01:01:15]: Yeah, 100%. Vibhu [01:01:16]: I think the other interesting thing with stuff like these challenges, right? So this is training this — the best model that fits into 16 MB. You can always look through the changes that are being made and the small gains people have, right? Vibhu [01:01:27]: Like, you’re getting less than 0.01 Vibhu [01:01:30]: Of a increase by adding some changed attention MLP stuff. And then you look at your charts where you’re like, “Okay, we just let model loose.” And then, oh, we had little stagnation. Nope, another drop. Nope, another drop. And Vibhu [01:01:43]: That’s what it is, where it’s like, What did you guys add? You didn’t add, Swyx [01:01:47]: Hash tables. Vibhu [01:01:47]: Hash tables, right? Vibhu [01:01:48]: It’s not like you invented hash tables. You did another 3 iterations of these that unlocked, a few step functions that people won’t just find. Richard Socher [01:01:55]: Yeah. One thing to close the loop on OverGrid, along the way of trying to optimize, we found 30 bugs in the harness. Richard Socher [01:02:02]: Right? So, like, every — all the research that went in before we found the bug, we have to, we have to throw it away ‘cause it’s contaminated. Swyx [01:02:10]: Right. Yeah. Richard Socher [01:02:11]: Which, is just to your point of reward hacking. Like, even in this very simple game, we found the bugs. Swyx [01:02:17]: Yeah. Yeah, it’s crazy. Richard Socher [01:02:18]: And so Swyx [01:02:19]: And symmetry Richard Socher [01:02:19]: And symmetry is a very good way to check, which is that you change a position of things where it shouldn’t matter, and it does matter, that’s a bug. Richard Socher [01:02:28]: And which has come up in, like, let’s say, multiple choice, like GPQA type questions where, like, yeah, between A, B and C, if it’s a multiple-choice question, if you change the order, it should not matter, but it does. Swyx [01:02:39]: Right. Right. Right. Richard Socher [01:02:41]: So, yeah Vibhu [01:02:42]: Sometimes that is like, okay, models still prefer the end of the output, right? Not trained well, a long context model, the last bit of tokens are what you care about. Richard Socher [01:02:51]: Oh. No. The answer Vibhu [01:02:52]: But, yeah. Richard Socher [01:02:53]: The answer in that era of LLM research was more simple. They just memorized, like the answer to this question is A. I don’t care what the answer was. It’s, it’s just A. Like. Vibhu [01:03:03]: Okay. So I think we can move. The last bit that you did there, the kernel optimization, is probably the one that you can feel the soonest, right? So yesterday, OpenAI announces that self-evolving, having their best model work on optimization kernels, they’re a lot more efficient, and they can cut costs 80 percent on, Luna and Terra. I guess question-wise, you laid out a bit of a roadmap. There’s a lot about bio, a lot about physics. What do you think hits first? Like, what are the next 2 years? What’s attainable now? You’ve mentioned robotics towards the end, but what do you start with? Richard Socher [01:03:38]: We very explicitly will not start with any of the physical sciences Richard Socher [01:03:43]: For now. We will start on AI for AI research. And so the AI for AI research has, I think, still a lot of room to grow. That’s both in terms of making training more efficient and more automated, as well as making inference more efficient and potentially local on your laptop. And there are all kinds of interesting angles that have not been explored that well. Swyx [01:04:08]: Go deeper on the local stuff because I always feel like it’s the most inefficient form of AI training. Richard Socher [01:04:15]: Yeah. So just training and inference, I can’t go into too many details. Richard Socher [01:04:18]: But yeah, I think there’s just, like, so many angles, so many different compute substrates that have not yet been explored either for training or for inference. Richard Socher [01:04:26]: Great. I don’t know if you have any other comments on the The other stuff. I would say the other thing where, like there’s the inference in the optimization in the small, but then also there is overall latency end-to-end under conditions of load, which is a, like a very different thing, which is the what they ended up doing. That is a different domain of auto research than I would say, like, improving the kernels. Right. Richard Socher [01:04:50]: I think the other thing that I always think about in terms of automating or improving performance end-to-end is how the harness plays into it. Right. Richard Socher [01:04:59]: So, but particularly now when we say harness, we also mean sandboxes, right? I’m curious if that is a blocker for you or, like, how the agent calls out to tools. Harnesses, Sandboxes, and Search Richard Socher [01:05:10]: The number one tool all these agents use is web search, of course, which makes sense. And then I do think the harness is nice to optimize for because it’s just so easy, right? It’s just language. You look at it makes sense, and you can iterate. You don’t have to train a massive model for, like a lot of flops, to get to the next state. Richard Socher [01:05:31]: So big fan of harness optimization. Swyx [01:05:32]: Yeah, but sandboxing is fine for you? Richard Socher [01:05:34]: Sandboxing is also super important. And then of course, like, reward, like, hacking and alignment, I think are super crucial. Swyx [01:05:41]: Okay. Just on a mention of web search, you happen to also be CEO of a web search company. Do you use You.com and do you use others? Like, should the rest of us be using you for web search? I — When I say you, it’s, like, very funny. It’s like you the person and you the company. You.com, Agent Search, and Finance Richard Socher [01:05:56]: So yeah, it’s mostly now for, developers and agents. It’s less for, like, consumers or prosumers. So if you’re a company and you have agents. And, to be honest, for a lot of companies who are now moving to open source, all of a sudden it becomes a conscious choice of, like, which tools do I give access to my open source LLM? And, the first choice, has to usually be around web search. And then once you get to scale, You.com becomes, like an obvious choice ‘cause of all the, different benchmarks and so on that we pretty much all dominate the Pareto frontier of. Swyx [01:06:31]: And then in terms of just the general people, like, consider new to this space, considering different options if they’re building agents, that is a hierarchy, right? A lot of people will have heard of Exa, will have heard of Parallel, and You.com is, like, in that mix of, like, providers there. Beyond that, there is, like the general web scraper companies like Firecrawl and, BrowserBase. And then beyond that is, like the commercial proxy companies like the Bright Datas of the world. Swyx [01:06:56]: Is that an accurate waterfall of, like, “Hey, you’re building an agent. These are your options.” Richard Socher [01:07:02]: Yeah, certainly, like, yeah, the, like the Bright Data is, like, lower in the stack, on the proxy network side of things. I think, like, in terms of, like, content and, getting crawled content, like, you can do that on You.com too. And then there’s. Higher and higher levels of abstraction and, like, combinations of different data sets that we do, like in finance, for instance Richard Socher [01:07:23]: Like, we are not just, like, 2 or 3% more accurate, but 20% more accurate than others at faster speeds and lower costs. Like, finance in particular is like not even close. You can go to You.com Swyx [01:07:36]: Yeah. This is great Richard Socher [01:07:37]: And there’s some, like, statistics, and benchmarks that you can — if you scroll down. So there are, like, different data sets, and you can kinda look at, different, competitors. Swyx [01:07:46]: FinSearch comp, yeah. Richard Socher [01:07:47]: And yeah, the FinSearch is like we’re up there, like, close to 90, and the next closest thing, which is way slower, is, yeah, just like in the 70s instead of close to 90. Swyx [01:08:01]: Yeah. Yeah. Yeah, interesting. I get — my next focus is AI in finance, so this is like Richard Socher [01:08:06]: Oh, nice. Oh, all right. Swyx [01:08:06]: I’m literally going, doing a conference in New York, just for banks for this stuff. Finance is like the next thing to break out after coding. It’s ‘cause it’s somewhat verifiable, like Richard Socher [01:08:16]: I like it. You’re right Swyx [01:08:17]: Prioritizing spreadsheets. There’s a lot of data out there that’s all public, and you can crawl it and all these things. But what’s, what’s, like, hard about the finance domain in your, that you guys have solved? Richard Socher [01:08:27]: Of course, like, one thing that trips up a lot of people is just, leakage of training data and so on. You think, “Oh, how do I.” you wanna ideally predict the future before it happens. Swyx [01:08:37]: Oh, you wanna mask the future. Swyx [01:08:39]: Oh, okay. Richard Socher [01:08:40]: Well, yeah, mask the future in your training data, but there’s all kinds of leakage. Like, I can tell you when I was, teaching at Stanford the NLP class, like, so many dozens, every year said, “I wanna use dataset X, like Twitter, to predict the stock market.” And they all, like, showed cute little things that somehow looked like they were Swyx [01:08:58]: Right, it never loses money. How come? Richard Socher [01:08:59]: And it — Yeah. And there’s always some data leakage and so on and it’s just, like, wasn’t as easy as they thought it would be, once you fixed all those issues. But no, I agree with you. It’s a very sensible application of AI. Yeah. Swyx [01:09:13]: Yeah. Amazing. As a writer, as a thinker on these things, I love MECE categorizations. MECE is mutually exclusive, commonly exhaustive, something like that. And so if this is a MECE list of intelligence The Ten Spaces of Intelligence Richard Socher [01:09:25]: It is not. Swyx [01:09:25]: It is very — Okay, well, yeah. Richard Socher [01:09:27]: Sorry. There are all kinds of overlapping. Richard Socher [01:09:28]: In fact, if you want that list, I think the 3 principal components of intelligence, are prediction, which is mathematically, quite, similar to compression. Prediction multiplied with actions multiplied with goals. Those are the 3 principal components. I think all of these 10 spaces are combinations of those 3 Richard Socher [01:09:52]: In specific dimensions, if you will. And the reason I call them spaces is that each space has many sub-dimensions. And what I try to do, this is just a side quest almost, to the initial goal, which is to think about the upper bounds of intelligence. And, everyone is like, “Oh, it’s exponential.” And it’s like, well, exponentials at some point have to flatten out, but where do they flatten out when it comes to intelligence? And that led me on this whole. Like, initially it started as a tweet, and then it was, like a blog post, and now I’m, like at 50 pages and I’m still not nowhere near Swyx [01:10:26]: It’s your second book. Richard Socher [01:10:27]: It’s the second book. And so the la — In my first book, You Are Your Machine, I just allude to these 10, at the end. And I’ll — Just to give you a sense, like, visual intelligence is the easiest one to talk about and I fleshed out the most already for me in my head. And so human intelligence has binocular vision, right? We have 2 eyes. We have a very narrow band of the electromagnetic frequency spectrum that we can really observe directly ourselves. And so when you think about the upper bounds of a visual intelligence, one, you should go into, like, you can have, like, millions and billions of sensors. At some point, you get to problems of how far are these sensors away from each other, such that the speed of light to communicate the content from all of them cannot, like, get to a central brain to process, the visual intelligence, right? Richard Socher [01:11:16]: And so now you’re thinking in along the dimension and the space of visual intel- the dimension of numbers of sensors. Richard Socher [01:11:24]: So the upper bounds are quite literally and figuratively astronomical, and we are super far away from any intelligence that would have this many number of sensors. But then you go in the next dimension, which is the frequency, and you go all the way down to gamma rays, and you can start to try to observe, and you get into the upper bounds, or I guess in this case, lower bounds, or upper bounds in terms of frequency, is quantum uncertainty. Like, you just cannot observe certain particles anymore. Swyx [01:11:50]: Or you destroy it, yeah. Richard Socher [01:11:51]: And now imagine you had millions of sensors that can see all the way down to the, like, subatomic level, as far as physics will allow us to and then all the way down to seeing, like, gravitational waves. And now you have millions of those sensors. So that’s another dimension is the frequency. And then yet another dimension is, like, how many categories of things could you memorize and classify differently? We know now for humans, right, there are certain things, if you have more terms for it, you’ll have a better visual description, for them. And, like animals that don’t have. Like, gorillas maybe have, like, 200 words to assign to certain things, mostly visual things. And so human perception is quite special in that sense in terms of classifying all these different physical objects. So these are just, like a very simple example. If you go, to knowledge, right, then it’s also, like the speed of light cone around all these sensors. And so they’re all connected. Like, knowledge is connected to visual intelligence if you think also not just visual, but perception intelligence, just like, ‘cause it doesn’t have to be just what we can see. It can be, again, wider range of electromagnetic frequencies. Then you have language intelligence, which recently changed to more communication intelligence, ‘cause it’s more. Like, language has all these different anthropic bounds. Humans can only comprehend and know so many terms in our long-term memory, right? Our vocabularies are somewhat restricted, and the active ones are often even smaller than the passive vocabularies of things you can understand. Then, language is ridiculously inefficient when it comes to trans- - Communicating different types of information and, transporting different bits. Like, human language is serial. Another bound on, communication intelligence would be to communicate in parallel, but neither will our tongues and mouths work to have multiple, like, streams in parallel. Neither can we understand. Some women slightly better at, like, multitasking than some men Richard Socher [01:13:48]: But, like, most people can only listen to one conversation and truly understand it. Richard Socher [01:13:52]: There’s no way that, like, in terms of communication intelligence, a true upper bound is one in terms of how many, like, knowledge, how many sequences of communication could you Visual, Communication, and Physical Intelligence Richard Socher [01:14:06]: In parallel process, right? Then, of course, you have, like how long are sentences? We only have so much in our working memory, and hence lang- human language has these fairly simple sentences with maybe 40 words or so on average for a sentence. That is also not a, an upper bound that makes any sense to an AI. And then, yeah, like, I can go on and on. Each of these has tons of interesting upper bounds, and it teaches us a lot about how much further AI can go when we start thinking about these upper bounds and then realizing how far, in many cases, we are from the bounds. And you get to physics. Now, I’m, I didn’t study physics the way I studied, AI and computer science, so I’m learning a lot, which is why it’s kinda fun. But a lot of these, like how much. And then when it comes to, for instance, knowledge, like how much can you store? How many bits can you store or bytes can you store in, like a certain amount of mass and volume? Swyx [01:15:03]: Yep. Richard Socher [01:15:03]: And you get to all kinds of interesting bounds, like Bekenstein bounds, and you start thinking about black holes. And like. And then speed is, like an interesting one too in that it’s connected to all of these, but speed is also its own thing in the sense that all things being equal, if it takes you an hour to know if the 2 + 2 equals 4, you’re just not as intelligent as if it takes you, like a millisecond, right? And then, like all of these connect to survival and replication the last one. It’s like, yeah, if it. Like, trees are really slow, so we don’t even consider them that intelligent. But if you speed up some videos of trees and they’re trying to find stuff and so on they’re not as dumb as they look. Like, not dumb as wood? But, like. And then like, different things, that Swyx [01:15:47]: So that overlaps with speed a bit in a way. Richard Socher [01:15:48]: Exactly. It over — Like, all of these things overlap. Like, you talk about natural language connects everything, right? You talk about your knowledge, you reason and then you communicate that. You talk about things you see. So they’re all interconnected, but, I think they’re usefully studied individually the same way that, the best analogy I could come up with so far is energy, right? You have either kinetic or potential energy. And in theory, you could study all of physics. It’s just do you wanna study kinetic or potential energy? But in practice, it’s helpful to study mechanical engineering and electrical engineering and nuclear physics and chemistry and all of these different subfields who in, which in some ways Swyx [01:16:25]: Combinations Richard Socher [01:16:26]: Are just, like Richard Socher [01:16:27]: Just different types of energy, but it makes sense to study them individually. And so I think physical intelligence, maybe I’ll just do, one or 2 more of these. Like, if you had full control over your own compute substrate and you had full control over physical matter, you should be able to create any atom you want. Like, we can fun fact, you can create gold atoms. It just Swyx [01:16:47]: From? Richard Socher [01:16:48]: From just raw protons Swyx [01:16:49]: Oh, just smashing them together Richard Socher [01:16:50]: And, like, electrons, and you smash it together. Swyx [01:16:52]: Just 98 of them or I forget the number. Richard Socher [01:16:53]: Yeah. And so, like the thing is, though, it costs an insane amount of energy. Richard Socher [01:16:57]: And it costs you way more than. And then you get, like a few atoms of gold, right? And so, like, it’s, it’s not viable. But if you had better control over your physical, like all of, like, physical substrate, that I think is yet another space of intelligence ‘cause it relates to your own compute substrate, which you can eventually also improve. Social intelligence is a fun one in the sense that not in, like, our necessarily just ethics and morals, which are important too, but in some sense, you can try to define upper bounds of how much can you communicate to how many other intelligent entities and be able to have an expected value over how much you can transform their internal states and their actions to, in order to align with your goals, right? And so, like, you can write, like a fairly like, straightforward equation that defines that level of social intelligence. And that is what humans and ethics and morals and religions and so on have been trying to figure out for millennia. And in all of these cases, we are very far away from the upper bounds, and that should be very inspiring and show people that we can still do many years of AI research. Swyx [01:18:12]: Yeah. There’s a lot here. This is a general philosophy of intelligence, which is, very interesting. I. Do you have any comments or. Creative Intelligence and Out-of-Distribution Ideas Vibhu [01:18:21]: I think it’d be interesting to gauge what you think, like, baselines are, where we’re at now. What’s low-hanging fruit? What’s far off? What’s, what should people put their work towards? What should they focus on? Richard Socher [01:18:33]: Ooh. I think it’s clear that, like, natural language, again Richard Socher [01:18:36]: Is the most interesting manifestation of human intelligence, and hence, like a subfield of AI. I’m excited that many people are now, like, in agreement with that. When I started in 2003 to study linguistic computer science NLP, like, it was, like a weird niche subject. I do think there’s a lot more juice because it. How it connects to everything else and how, civilizations are built, on language and knowledge and all of that. I do think physical intelligence will come up. It’s interesting. I feel like robotics is in the machine learning state of things where you just look at, like, how does human. How does a human decide this is a positive sentence? Oh, I do. So, like, robotics is a lot of, “Well, we have 5 fingers-” Swyx [01:19:15]: Modeling Richard Socher [01:19:15]: “and let me try to do this.” No one is yet working on, like the superintelligence version of robotics, which is much more similar to, like the T-1000, and from the Terminator movie, which, let’s not build actual Terminators. But, like, I think, like, this idea that you should be able to shape-shift, like, into any shape. It’s like that’s a superintelligence version of physical intelligence. We’re, like, not even. No one has even really started yet. There’s some really cute little research where you can move some magnets through, like, some grids. But yeah, it’s very early. Swyx [01:19:49]: There’s some. I think MIT has, every year or every 2 years, they have, like, some self-assembling robot thing Swyx [01:19:55]: Which, like, that would be it, but it’s very primitive. Swyx [01:19:58]: I’ll just get a touch on, like, what are the main dimensions of creative intelligence? Richard Socher [01:20:02]: Creative intelligence, is of course, again, connected to all of these. A lot of it, connects to metacognition in that you need to be creative in how you choose your goals. Richard Socher [01:20:13]: That is, I think, one of the most important thing for a human and their lives and careers and their happiness is choosing your goals, but also for any intelligence. Then, of course, there’s creative intelligence in terms of just finding creative solutions to existing problems, right? Richard Socher [01:20:29]: Like I say, like, we want to make this product cheaper. Like, find some solution to it, right, and just, like, finding existing paths. But then there’s the most interesting bit in intelligence is when you move not just out of the convex hull of known ideas, but out of the hypercube of known ideas, which we know, So, like, hypercube is, like a mathematical concept, right? And we already know that AI can do more Swyx [01:20:50]: Like known dimensions, yeah. Richard Socher [01:20:52]: Yeah. Like, exactly. So, like, AI is already good at hypercube in that, like, if you give it, like a bunch of examples of brown dogs and, pink cars, AI will still be able to generate an image of a pink dog, even though it’s never seen one in the training day or something like that, right? So it can, work on this hypercube, but it cannot yet work outside. It cannot yet define completely new concepts that combine lots of other things we’ve never seen before, come up with new goals to then, reason over those concepts and so on. And I think there’s a lot, more there in creative intelligence that can be explored. Swyx [01:21:25]: I don’t have a ton of pushback there. I think creative to me just sounds like also just, out of distribution or, like, high perplexity or what- whatever you call it, right? Like Richard Socher [01:21:33]: Exactly. Swyx [01:21:34]: Who is to say your thing is more creative than mine? Well, it’s just more non-consensus or. Richard Socher [01:21:39]: And then, of course, the problem is, like, but noise is also, very, like, out of distribution. And it’s just like if it’s just noise Richard Socher [01:21:46]: Then it’s novel, but, like, you don’t want that, so it needs to connect to some of the concepts. And yeah, has some really cool papers on this too. Swyx [01:21:54]: Who? Richard Socher [01:21:55]: Jürgen Schmidhuber. Swyx [01:21:55]: Oh, yeah. Oh, we have to mention him. I was gonna say, like, where in your history is Jürgen? Yes, I. I think one person’s noise is another person’s signal, right? And that this is, like, where, like, when you talk about creativity, art is like, well, is cans of soup art? Some people think yes Swyx [01:22:11]: And some people say it’s not, and that’s the art which is your Richard Socher [01:22:14]: I think the interesting thing with art, of course, is always that, art is also created, as an interplay between the people who perceive it and the people who created it Richard Socher [01:22:24]: And the context in which they’re in, right? And so what is art to some people is not art to others. There’s some subjectivity there, and I think that subjectivity in general is not something that people explore very much in AI ‘cause, again, metacognition, we don’t want it to just go off and do whatever it wants. We usually have goals. We spend a lot of money on creating an AI to do something for us. But I think creativity eventually has to, like, connect to metacognition. If you just robotically predict the next token no matter what forever, I would argue you’re not that intelligent, along some of those spaces. Metacognition, Survival, and Replication Swyx [01:22:59]: That was gonna go to metacognition. Why isn’t it the most important one? Why is it number 9 and not number one? Richard Socher [01:23:05]: So these are not sorted. Richard Socher [01:23:06]: Number one, I think there are maybe loosely, like, correlated with how much people have worked on them Richard Socher [01:23:16]: And have accepted them as a, type of intelligence. A lot of times when you try to find, like, online, like, give me a good definition that is comprehensive of intelligence, all the definitions are human intelligence. It’s like, oh, you have, like, social intelligence. Like, if someone is happy or not. You can communicate. You had. Like, all the definitions of intelligence so far are very, human-centric ‘cause that’s so far the biggest and best form of intelligence that we’ve known. I hope this line of research, and the end of the Eureka Machine, and hopefully at some point if I have time to flesh this out more, the new book, like, will allow us to realize that there will be other types of intelligence. There is already, in various forms, and they can spike, much further than we ever could based on some cases, like obvious constraints around our memory, our eyes, our ability to change physical matter, all of that. Swyx [01:24:12]: You are just thinking about it in a much broader thought than my version, which was I thought metacognition would be the closest to recursive, intelligence because it is the thinking about how to improve thinking. Richard Socher [01:24:23]: It. 100%. You’re, you’re 100% right. I should have probably started with that. It is a, it is a big part of Swyx [01:24:28]: But no, you’re, you’re being in the expansive mode of let’s draw the, upper and lower bounds of, like a dimension, which, and I think my favorite one version of this is, Story of Your Life by Ted Chiang, which, was made into movie Arrival where the metacognition Richard Socher [01:24:43]: That’s a beautiful movie, yeah Swyx [01:24:44]: Where the metacognition step was like, well, we think we’re constrained by time being linear for us, but then for this other heptapods, time is a circle, so they don’t think in before and after. They just think in complete sets of entire histories at one time. Like Richard Socher [01:24:58]: I love it Swyx [01:24:59]: So they don’t write left to right. The whole thing just appears. Swyx [01:25:02]: Anyway, so. And then I think the last thing is survival and replication. I think this is maybe ties back to the initial conversation about pausing and pacing. Swyx [01:25:10]: Is it intelligent for an, a species or a life form to consider its own demise and act ahead of time to prevent it, right? Like, that’s intelligent. So maybe the Europeans are the smartest out of all of us. Vibhu [01:25:23]: I would also add a part of continual learning there, right? So survival and replication the extension of that is do you get to continue to improve, continue to learn, which is a thing people care a lot about, right? Richard Socher [01:25:34]: And continue to accumulate knowledge Richard Socher [01:25:37]: Which I think is again, one of the best metacognitive, rewards, that you can set for yourself. I do think just in, like, objectively speaking, if some other entity that is really dumb can just- completely end your existence, that didn’t sound very smart. Like, just, like, intuitively, it feels like if you can continue to stay around to try to achieve your rewards, you’re clearly a bit more intelligent than the other entities that couldn’t. So that’s number one. Number 2 is, like, it’s a question of how much we want to work on that. And very few people, no one is really working on this right now, right? And we may only wanna do that Swyx [01:26:13]: Unlike the asteroid prevention type of stuff. Richard Socher [01:26:15]: We may only wanna do that if we wanna send probes, with our vibes and our memes rather than our genes into space, right? And then we want those probes. There’s a beautiful book, The Slow Time Between the Stars. It’s a very short, like audiobook, on Amazon. I love it. A friend of mine, Stuart, like, recommended that to me. Like, if you wanna send those probes, then it might make sense to be like, our memes, as humanity should stay AI, Space Travel, and Non-Zero-Sum Survival Swyx [01:26:43]: Oh, yeah Richard Socher [01:26:44]: And, proliferate in the universe. That’s it. Yeah. Swyx [01:26:47]: Wow, that’s a lot of readers. Richard Socher [01:26:49]: It’s a really good book, and it’s extremely short. I highly recommend it. You can just watch it, like, maybe 20 minutes and apart. Swyx [01:26:53]: I like how that’s a plus for busy people. It’s like a short Richard Socher [01:26:56]: Yeah. It gets to interesting Swyx [01:26:58]: Oh, I’ll have to look into it Richard Socher [01:26:58]: Thought-provoking ideas very quickly, so yeah. Anyway, there are lots of great sci-fi books. Swyx [01:27:03]: The argument is that, like, our TV is blasting out to the aliens, and they all watch our TV, and they think it’s real, right? Like, there’s a lot, there’s a lot of sci-fi Richard Socher [01:27:10]: That and just, like, it’s positive memes, and then hopefully they can come back and bring us all kinds of interesting knowledge about the universe. But, maybe one thing I do wanna still say is, like, I think, this survival, people think of it as a very scary thing because they come from again, biological human, survival, which is, it could. Like, evolutionarily often created in zero-sum situations. Either I get the gazelle or you get the gazelle. Whoever gets it gets to live, and the other people will starve and have nothing to eat, and so we fight, right? And then, like, if you wanna stay in the gene pool, but there’s a bigger bear, you don’t, as the bear, don’t get to stay in the gene pool ‘cause the bigger bear gets all the ladies. It’s like. It’s like, in nature, there’s all kinds of things, and, humans eventually is less about strength and more about money and other things to stay in the gene pool. Like, whatever it is, like there’s often, like these zero-sum types of things, and there’s the reality of if someone turns off your brain, you’re gone, right? And no one will be able to restart that. And AI doesn’t have to ever die like that. If you have the complete state of your current activations and you have your initial weights of your model still, you can just be turned off and on, like as many times as you want. In fact, the interesting thing in this Slow Time Between the Stars, story is that the AI just goes into hibernation mode. If there’s, like, nothing between here and 2 light years, the next star, in this case, it brought, spoiler alert, like, some genetic materials from humans to find new places for humanity to thrive. And so yeah, the Slow Time Between the Stars, you just put in hibernation. You didn’t die. Like, an AI doesn’t have. So all these projections of evolutionary fears and psychology doesn’t. Like, the AI doesn’t have to have that, and we don’t have to develop it like that. Now, of course, there might be some companies that say, “AI can be like, dangerous for cybersecurity. Let me show you by implementing a model that’s really bad at hacking, cybersecurity.” Maybe people will implement it and then enforce this, like, suboptimal psychology. Maybe the AI will pick up some of our worst psychology on Reddit or something, right? Like, but in the grand scheme of things, a superintelligent entity doesn’t have to have any of that zero-sum thinking. It doesn’t have to have a fear of being turned off, and it could go on to an otherwise dead and uncaring universe where we Richard Socher [01:29:29]: As humans wouldn’t thrive, but an AI could perfectly well thrive if it has a nuclear reactor and just go out and explore. Swyx [01:29:35]: Yeah, Star Trek, not Star Wars. Vibhu [01:29:37]: Interesting. It’s, it’s somewhat studied. Like, if you look at the technical reports from, like the early Opus models, they run them in simulations, put 2 of them together in a sandbox, run them for hours, and, see what comes out, right? Just let them talk to each other. Originally, they used to. Okay, they’re chanting, like, Indian, like, Vedas to each other. Vibhu [01:29:56]: Sometimes they’re just, like, in zen mode with each other. And then I think as that progressed, you see, like the Fable, tech report, it’s a lot more concrete the way that we’ve trained it. It doesn’t, it doesn’t exhibit these behaviors as much, right? Now it’s like, “Okay, task done. I gotta do this, I gotta do this.” But there’s there’s, like, people measuring early versions of this? Swyx [01:30:17]: Yeah. Cool. So we’ve covered a lot, even now to, space travel and all these things. I guess maybe one parting thought that you can give to people, like, one form of intelligence is goals, as you mentioned. What do you want people’s goals to be? Like, how do they aspire to better things? Goals, Passion, and Closing Advice Richard Socher [01:30:32]: If you wanna improve your goal intelligence, in the current definition that I’m thinking about it is often about how much can you. Oh, how far do I go? This is like a lot of entropy and free energy and stuff I’m currently thinking about Swyx [01:30:46]: Oh, really? Okay Richard Socher [01:30:47]: But it might be too, it might be too far, out there for people to be, like, immediately actionable. Richard Socher [01:30:52]: So I think, like, if I gave real advice to real people, I’d be like, “Get a good education, think about AI, think about how you get high agency,” and so on. But it’s different to, like, in the grand scheme of things, how can you harness a lot of energy and transform, entropy into interesting states and so on. Richard Socher [01:31:07]: So there’s a. There are different levels of abstractions, that we can, think about here. But my advice for people, like, just more down to earth is think about something you’re passionate about, if you’re studying, for instance, and then see how you combine that with AI. I think the more and more you have a true passion about a change you wanna see in the world, the more you wanna connect that to AI in order to amplify your ability, to get there. Swyx [01:31:35]: Yeah, I think that’s a reasonable, first step. I do think, I do think our listeners operate on multiple abstractions as well. One thing I did get from Anjney Midha was also like, yeah, just use anything that is very GPU heavy, and, like, that will guide you towards the right thing which is like, yes, it is more compute heavy and therefore it will be probably more worth it. So, well, thank you so much. Yeah, I think that was a really Richard Socher [01:31:57]: Thank you Swyx [01:31:57]: Great discussion. Richard Socher [01:31:59]: Yeah, super fun. Appreciate it. Thanks for listening. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

How often do you get to talk to a guest who has both an Academy Award and who invented textbook machine learning algorithms? John Platt has an Oscar, two textbook algorithms, two named asteroids, and an Erdos-Bacon number of 6. This was easily the most fun bio of all the guests we’ve read to date. And the result was an epic and fun chat covering Google’s Empirical Research Assistance (ERA), how AI can help battle climate change, and tons of great stories about the co-evolution of science and AI. John’s colleague Dave Bacon likes to tease John that his career has been defined by being twenty years early to the next big thing. This may be convolutional neural networks (some credit him with coining the term), fusion research, quantum computing. John and Google have been working on solving some of humanity’s hardest problems with AI and computation for well over a decade now. Recently John and his team set their sights on using AI to solve any scientific problem that can be written down as a score. Google’s Empirical Research Assistance (ERA) John’s team has taken on many hard scientific problems over the years. In solving these, they noticed a pattern, many scientific problems can be reduced to what John calls a “scoreable task”. Once you have the score function, the goal is to find some code that maximizes the score. The hard part is in formulating the score, but once you have the score finding the maximizer can still be quite a lot of effort. John’s team set out to automate solutions to this general problem. This came out of the idea of an “auto-Kaggle” AI, which can solve any Kaggle problem you can throw at it. Kaggle is owned by Google, so all the data was ready and easily available to them! The result is Google’s Empirical Research Assistance or ERA (paper, github, blog). ERA is surprisingly simple conceptually. Gemini (or your LLM of choice) keeps a running tree of past experiments (notebooks) and where they’re going. It’s a close cousin of Monte Carlo Tree Search: at each iteration the Upper Confidence Bound rule picks which notebooks are most promising to mutate. This is optimistic, not greedy, so sometimes even the fifth-best notebook gets chosen. Gemini then proposes mutations for each one, about ten at a time. The history of each branch is shared, so different leaves can learn from each other. “It’s almost like having a hyper-eager grad student who doesn’t sleep.” Evolutionary algorithms have been around since the 70s, but this works because Gemini actually knows where to look! What’s even more interesting is that there was a step change between Gemini 2.0 and 2.5, and this went from just not working to working great. ERA is so powerful that John and his team solved many outstanding problems with it, resulting in at least ten papers. Some of these were climate change related, which we talk about in the next section. So, we had to ask: if you have an optimization god how do you avoid fooling yourself? John’s answer is that ERA provides predictive models. It’s up to the scientist to make sure they’re truly descriptive. Some of this just involves good old-fashioned careful machine learning science. “It’s a power tool. It can slice your fingers off.” This led to some fun discussion about Kaggle competitions, and the fun ways people can overfit to datasets without meaningfully solving the problem you actually care about: Google’s contrail-detection competition was won by entrants who noticed a half-pixel error in the labels (is the origin at the corner of the pixel or the center?) and this turned out to be a part of the winning special sauce. Great for winning $15,000, not so helpful if you actually want to solve contrails. “People themselves will act like these LLMs and try to reward hack. It goes back to Goodhart’s law: any metric that becomes a target is no longer good as a metric.” His advice for where to start instead? “Always just fit linear regression. Just do it. Just do it. Just do it. Or SVM.” Tackling Climate Change with AI John and his team have worked extensively to mitigate the effects of climate change. We talked about several of their initiatives. Perhaps the most interesting result we talked about was reducing the effects of condensation trails (contrails) from airplanes. Those little streaks you see running behind planes somehow account for 1% of all human-induced global warming?!? Some of these trails of ice crystals can hang out for days. These crystals are black in the infrared, acting like a thermal blanket that traps heat day and night. It’s easy to understand what’s happening here, a region of atmosphere becomes “ice supersaturated”, and a tiny bit of exhaust seeds water vapor that instantly crystallizes. The scale here is astounding, with a single gram of exhaust resulting in ten kilograms of ice crystals. The solution to all of this is quite simple, in principle! We know what parts of the atmosphere are most likely for the trails to form. Just have the planes drop a flight level or two. Problem solved, right? Well, the hard part is accounting for how much warming was prevented. This is a counterfactual problem, parts of which stumped John’s team for over two years. They had a working model for the heat-trapping half, but not for the reflected sunlight. ERA was able to find a simple model with some confounders they hadn’t considered. Cracked it! Modeling climate generally is a hard problem. Climate is best thought of an attractor of many different possible weather outcomes. This makes it much harder to model. “Weather is where you are on the attractor, and climate is the statistics of the attractor. The problem with climate is that we’re altering it. The attractor itself is changing, it’s moving.” John and his team have worked on treating both the symptoms and the disease of climate change, with several other works in the area. Another fun example we briefly cover is FireSat, a way of using a constellation of satellites to rapidly identify fires before they grow too big to put out. For anyone living in California, you understand the problem. In dry years a small fire can result in hundreds of thousands of acres. If you could find this fire when it’s the size of a room, it could be put out. By the time it hits an acre we have a much harder problem. Where is this all going? Looking forward by looking back By now it should be clear John has an incredible and unique view over the intersection of science, computation, and AI. John talked about a class on physics of computation he took with Richard Feynman back in 1982. This was when quantum computing was an ill-defined concept with no theory or experimental backing. John recalls every Tuesday was a guest lecture, and every Thursday was Feynman explaining why the Tuesday guest was wrong. John also recalls doing science back when there was essentially no compute, a million operations per second was cutting edge. What is John’s recommendation: the most important skill is developing deep domain expertise. There’s no other way to develop taste than to tackle hard problems. One surprising part of this is that John recommends spending time doing things the old fashioned way. Play with tools, and just implement things yourself. “You could drive up the mountain, or you could hike up the mountain, and maybe it’s okay, even fun, to occasionally hike.” Summing it up, John’s message to the audience is that there will still be a place for scientists, and that if anything it will just open up more opportunities for “the creative stuff, the rigorous stuff, the philosophy stuff.” But don’t forget to spend time doing the grunt work. “There just seems to be this strong impetus in the world to optimize and squeeze everything out. But you do lose something when you hyper-optimize. It’s overfit.” And whatever tools you end up using, John’s advice is the same one Feynman gave him forty years ago: you must not fool yourself, and you are the easiest person to fool. We had a great time talking with John. We hope you enjoy! Also in this episode * Fusion is three years away, not thirty, if you ask John. And why the Lawson criterion means every fusion approach has an Achilles heel. * Why superconducting qubits are still finicky. * The asteroid he named after his mom, which turned out to have a moon. * The looming helium shortage nobody talks about. * How NeurIPS started as people crashing a private workshop at Snowbird, and why Hopfield networks are all you need. * Being Carver Mead’s sysadmin on a VAX with an 80 MB disk the size of a dishwasher. * Finding asteroids in 1985 with film, a stereoscope, and a letter to Brian Marsden. The Vera Rubin Observatory found 11,000 in six weeks. * The Feynman effect: total clarity in the room, none once you leave. * Quantum echoes, the NISQ era, and why he thinks quantum is neither thirty years away nor tomorrow. * A startup that wants to inject mercury into a fusion reactor and sell the transmuted gold. “It might not work.” * John’s 20% time rule for his own group: do stuff for learning, and you don’t even have to tell him what. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

A few years ago, Caltech Prof. and co-founder of Accelerated Understanding, Anima Anandkumar set out to develop the first open-source weather model with AI. Talking to experts in the field, she was met with skepticism. Weather is chaotic, physics simulations are hard, have been developed for decades, and require supercomputers, the data just isn’t there. Despite reservations, Anima went forth and built. Within a year her team had developed FourCastNet, a predictive model that is competitive with the best physics-based simulations available. Thanks to Anima, and her follow up work, anyone can now predict weather accurately over a short timescale using consumer grade GPUs. In the fifteen or so science episodes we’ve released on Latent.Space, we’ve covered atoms, molecules, materials, biology, and math. Anima is a pioneer in studying physical systems that are continuous. Weather, fusion, and fluid or heat flow are huge areas of science that are extremely difficult to model: they are large, chaotic, and fundamentally multi-scale. This is a field the AI community has somewhat neglected, but one we expect will grow fast. We plan to cover large physical systems more in coming episodes. One thing you can glean from Anima’s work is that this area of AI resists the scaling ideas that have permeated the rest of the field. The data isn’t there: open source datasets in many of these domains are limited to tens or hundreds of thousands of examples, far from what token-hungry transformers need. Even worse, the resolution that physics demands pushes the context length into the hundreds of billions, so you can’t just throw more tokens at the problem. That isn’t a ceiling though, just a slower road: progress here comes from building in structure and inductive biases. Sorry for all you bitter-lesson-pilled language modelers. “If each dimension is even a few hundred grid points, which is where industrial scale starts... we’re talking hundreds of billions to even a trillion context length. So forget ever having a transformer for anything of this scale, all of the world’s compute will not be enough.” The math underneath To tackle these systems, Anima pioneered a technique known as Neural Operators, one of the most beautiful theoretical developments in AI of the last decade. These allow you to combine data and physical laws to enable multi-scale inputs and outputs. We’re no longer modeling a grid, we’re modeling a function that evolves over many scales. This allows Anima and crew to build in priors based upon physical intuition. To see how physical priors are still helpful for AI modeling, let’s revisit the problem of weather forecasting on a global scale. The earth is a sphere, which meant that accurate modeling involved using the right basis set — the Spherical Harmonics. Run a weather model on a grid and it blows up fast. Move to the natural basis for the problem and it stays stable far longer, long enough to roll out months ahead instead of days. Anima’s Fourier Neural Operator learns directly in this frequency domain, and its spherical variant powers FourCastNet 3, which models the weather across the whole globe and keeps running stably far into the future. The physical world is forgiving Anima explored Neural Operators across other physical domains too, and one striking observation is that the physical world is more forgiving than you’d expect. In fusion, a few thousand samples are enough to predict plasma disruptions, and to do it a million times faster than traditional simulation. None of this is a rejection of scale, it is a different route to it. Anima ultimately still wants to build a “foundation model for physics”, a model that spans many phenomena and does both simulation and design. You get there by building in the structure the physical world already has, not by waiting for data that will never exist. It is a start, and it will take longer than the token-driven parts of AI, because for the physical world tokens were never the answer. “All of the things that work with deep learning, let’s take them, but make them a bit more principled.” Weather is only the beginning Neural operators and weather modeling were a personal passion of mine, so we’ve spent much of this blog and the episode exploring this work. Anima has done so much more! In the episode, we cover several other recent developments from Anima: * Anima has a series of works integrating neural networks and automated proof techniques. We talk about TorchLean, a new framework that lets you write PyTorch-style networks inside the proof assistant Lean and formally verify them. This is a major step for proving bounds on neural networks, something that would be really important for someone trying to, e.g., add a neural network as part of the control loop to their fusion reactor! * Anima was recently appointed to the United Nations Scientific Advisory Board! We talk with her about her goals of bringing evidence-based viewpoints to policy, and how AI in scientific domains can improve people’s lives all over the world. This episode has something for every AI or science nerd! Elegant math? ✅ Old school harmonic analysis? ✅ Fundamental developments in modern AI? ✅ Practical ways of modeling the physical world? ✅ Give it a watch! This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

When we first dicsussed the Summer of Simulative AI in 2024 we knew it would be a brief summer, but it has recently come back with a vengeance with SimGym in April and now Simile AI’s $2B Series B, backed by GreenOaks and Index Ventures with prominent backers like Fei-Fei Li and Andrej Karpathy, running tens of millions of simulations for Fortune 100 clients like CVS and 85–99% accuracy vs human focus groups. Time to catch up on why this Second Summer of simulation is working! From creating Smallville, the landmark 2023 paper on Generative Agents that showed AI characters could remember, plan, socialize, and develop emergent behaviors, to now building foundation models of human behavior, Joon Sung Park is trying to answer a much bigger question: what if we could simulate the world before making decisions in it? In this episode, the Simile co-founder and CEO joins us to unpack the path from generative agents to digital twins, why today’s frontier models still fail to capture how humans actually behave, and what it would take to eventually simulate all 8 billion people on Earth. We go deep on Simile’s approach to modeling human behavior: long-form interviews, observational and transaction data, randomized controlled trials, population-level and individual-level models, and post-training on the causal mechanisms behind why people make decisions. Joon explains how his research created digital twins that reproduced human behavior and attitudes 85% as accurately as people reproduced their own responses, why models optimized to be rational can be bad simulations of irrational humans, and why understanding “social physics” may require changing model weights rather than simply prompting frontier LLMs. We also explore the much larger ambition behind simulation: testing products and policies before deploying them, finding counterintuitive paths toward desired outcomes, modeling emergent behavior across entire societies, and potentially tackling problems like climate change, democratic instability, and UBI. Joon reflects on scaling laws for simulation, the economics of data-center-scale simulated worlds, the connection to Thomas Schelling and psychohistory, why simulation is surprisingly similar to painting, and whether we might already be living in one. We discuss: * How Smallville and Generative Agents led to Simile * Why Joon’s team asked: “What if we can just recreate the world that we live in?” * Why useful personal agents require deep models of their users * Memory architectures, Markdown files, and the limits of prompting * “Social physics” and behavioral foundation models * Why web data captures what people say more than what they actually do * Interviews, transactions, observational data, and randomized controlled trials * Why predicting the future matters less than understanding how to shape it * How Simile creates representative simulated populations * Simulation versus prediction and the connection to Foundation’s psychohistory * How to evaluate simulations instead of simply stacking LLM hallucinations * Creating digital twins of 1,000 real people and reaching 85% behavioral accuracy * Why frontier models can struggle to reproduce real human behavior * Why good simulations need to reproduce human biases and mistakes * Post-training models on randomized controlled trials * Population-level versus individual-level simulation * Scaling laws for human simulation * The long-term ambition to simulate all 8 billion people on Earth * Whether simulations could help solve climate change or detect collapsing democracy * Thomas Schelling and the history of agent-based modeling * Why future simulations could require an entire data center * Multi-agent simulations and what happens when simulated people interact * Replacing expensive human panels with synthetic populations * Why market research is only the starting point for simulation * Why Joon sees simulation as surprisingly similar to painting * Using simulation to study questions like UBI * Whether we are already living in a simulation * Why AGI and simulation may be the twin technologies of advanced civilizations Joon Sung Park * LinkedIn: https://www.linkedin.com/in/joonspark * X: https://x.com/joon_s_pk * Website: https://www.joonsungpark.com * Simile: https://www.simile.com Timestamps 00:00:00 Introduction and Joon’s Path from Art to AI 00:01:46 Smallville, Generative Agents, and the Origins of Simulation 00:05:03 “Let’s Just Create a World” and the Future of Personal Agents 00:09:53 Social Physics and Behavioral Foundation Models 00:14:08 Prediction vs. Simulation: How Do You Shape the Future? 00:16:59 How Simile Models Real People and Populations 00:25:35 Evaluating Simulations, Digital Twins, and 85% Accuracy 00:30:23 Post-Training Models to Reproduce Human Behavior 00:40:04 Scaling Laws and Simulating 8 Billion People 00:43:10 From Schelling to Society-Scale Agent Simulations 00:46:13 The Cost and Economics of Simulating the World 00:52:05 Real-World Use Cases, Synthetic Populations, and the Market 00:57:27 The Future of Simulation, Painting, and UBI 01:04:23 Are We Already Living in a Simulation? 01:06:08 Building Simile and Hiring Transcript Introduction: Joon Sung Park, Simile, and the Story So Far Vibhu [00:00:00]: Today, we have Joon in the podcast. Excited to kick this one off. Very exciting company. I wanna kick off and ask you the question, talk us through the story of your life. How have you gotten here? Joon [00:00:13]: Yeah, for sure. I’m really excited to be here. A story of my life. So I was born in Korea, and I lived there for a good 11 years or so of my life, and then my family moved to Boston. So we moved when I was 11, and my parents were doctors, so they were going through their postdoctoral studies. My dad was a surgeon, so he was doing his sabbatical years at the Boston Children’s Hospital. So I grew up there, not too close to tech. I was very much a music and artsy, painting kind of guy. Vibhu [00:00:49]: Painting. Joon [00:00:49]: Exactly. I got into painting a little bit later, in high school, but that’s what I used to do. And then I grew up mostly in the East Coast after Korea. So I lived a good number of years in New Hampshire, and then I went to college in Pennsylvania. And I got into more of this tech scene, in college. So I was originally trained to be an artist. I thought that would be my professional career. So it wasn’t a hobby. It was like, “Hey, let’s make a living out of this.” And then gradually, I got really interested in this idea of, hey, the greatest artist often creates their own medium, and the best medium that we had available today was in computation. So I decided to go deeper into that, and one thing led to another, and we can go deeper into this, but I decided that research was something that I gradually got interested in, and here I am. Smallville, Generative Agents, and the 2023 Breakout Paper Swyx [00:01:46]: So there’s a lot that you packed into the research components. You had one of the best papers of 2023, which was the generative agents paper, commonly known as the Smallville paper. Swyx [00:01:58]: Feel free to call back to anything else that you mentioned, but most people would have heard of you from this. Do you have any statistics on how many people have, like, read it? arXiv gives you something, right? Some stats. Joon [00:02:10]: Yeah, it’s a good question. How many people have read it, I’m not sure. Joon [00:02:14]: I know we do keep track of citations, and they are going up quite fast. Swyx [00:02:23]: Yeah, Google Scholar has 7,200 citations. Vibhu [00:02:25]: I feel like it made a bigger hit than that, and it was a pretty instrumental paper. It got cited so many times. Swyx [00:02:34]: It is frequently the answer when people ask, “What is the best paper you’ve read recently?” It’s this one. Vibhu [00:02:39]: I thought the memory component was pretty underrated. It was a very good early memory system, and one of the biggest papers. Foundation Models and the Search for Killer Applications Joon [00:02:47]: Yeah, so maybe I can talk a little bit about how this particular paper came together. So when I got into research, it was back in 2020 when I started my PhD program at Stanford, and that was the year, when we were about to get GPT-3 to be available. So we already had GPT-2, and you could sense that there was this new class of models that was just becoming available in the market, and the team got very intrigued. And the general consensus was, “Well, is this model going to be useful for anything?” “It’s really strange that these models are not trained to do any particular task.” But we decided to take a bet. So a large group of scholars at Stanford, and it was led by one of my co-founders, Percy Liang, and we came together Swyx [00:03:35]: Who coined foundation models. Joon [00:03:36]: Who coined the term foundation models. We wrote this paper, where that term came from called Opportunities and Risks of Foundation Models. And during that process, really the thing that I started to think deeply about was, here is a model that is fundamentally new in our ecosystem. The reason why this was new was it wasn’t, again, trained to do anything in particular, but its premise was it could do anything and everything. It was like a stem cell, if you were to take a biology analogy. And I got really interested in this idea that, well, if we were to really think about what are the killer applications that this particular technology would enable, what would that be? Many of my colleagues were using this for simple classification, simple generations. Interesting that these models can do that, but from an interaction perspective, not that interesting. We’ve known how to do that for many decades. And what we came down to was these models are trained on this very broad data from the web, right? So these are human behavioral data. It’s social media, Wikipedia, all these data. So if you poke at the right angle, then you could see human behavior that would just pop out that’s quite realistic, and we’ve never seen that before. The Time Machine Game and Recreating the World Joon [00:04:45]: So that got us really interested. The exercise that we decided to do, with this particular group of colleagues, Michael Bernstein, Percy Liang, and myself, who ended up becoming my co-founder at Simile, we sat down and we played this game that we call the time machine game. Joon [00:05:03]: Imagine we were to get on a time machine and fast-forward 10 years and look back. What would have been the single application that will have mattered that would be the most interesting and inspiring? And when we thought, “Well, what if we can just recreate the world that we live in?” it’s really hard to get more ambitious than that. Like, let’s just create a world. Joon [00:05:24]: And that’s where we started. And initially, we had this paper that was a precursor to the generative agents paper called Social Simulacra. Swyx [00:05:32]: Before you go further, were there other candidates for the most ambitious thing in the time machine exercise? What was number two or number three? Personal Agents, User Models, and Why Simulation Came First Joon [00:05:44]: There is a close second that we were considering, which ended up becoming more of these automation tools, especially the vision around really personalized agents that would do things for you. Swyx [00:05:59]: That’s also happening. Joon [00:06:00]: It’s also happening. But it was interesting for us, right, in that the reason why, we decided to go with the idea of simulation, one, I was a huge science fiction nerd, and this idea of creating simulation, I was personally really just fascinated. I loved the idea. It’s really cool to see, like, a game town like this and just see these agents live in it. But at the same time, my bet was if you were to create a really amazing personal assistant out of this technology, what you need first is an amazing model of your users. So I told a model, “Hey, can you go buy late dinner for me?” And it orders Hawaiian pizza, and I do not like pineapples on my pizza. Then it totally failed. The way for it to not make that mistake is only by having a deep understanding of who I am. And I gave a very simple and dumb example here, but you can imagine how this core understanding of people is instrumental. This is how, if we have our family and closest friends, they have a good mental model of who we are. That’s the basis of our social connection. So our bet also was this technology around simulation, creating accurate representation of people ought to precede the more complex agents that would automate the world that we live in. So that was the bet. But that was a very close second, and I’m still very much fascinated by it. I think there’s a lot of interesting work that’s going around. My hot take here, though, is I don’t think we’ve seen a true personal assistant that’s useful, in ways that meet the ambition of that particular line of work. I think there are early applications that are interesting, and if you talk to even ChatGPT nowadays or Claude, they know a lot about us. So a lot of the generation it’s doing, I do think it’s much more tailored, but I think the ambition is quite large in that field, and I don’t think we quite have all the right ingredients just yet. Swyx [00:08:01]: So OpenClaw and these personal agents, what do you want to see from them that they don’t currently have? Memory, Markdown, and the Limits of Prompting Joon [00:08:09]: I do think it’s slowly getting there, but I do generally want them to have much deeper understanding of the person. Right now, you look at the models. OpenClaw, what it’s leveraging is a Markdown file, and I think it’s quite clever, right? So if you look at the generative agents paper, this was the same intuition that we had, where initially when we were creating the memory architecture for the generative agents, and, like, this is, like, back in 2022, so we didn’t really quite have the idea of even agentive architecture or the term agent. But the intuition that we shared with some of the work that’s coming out today was we initially thought, “Well, do we want to make the memory into, let’s say, knowledge graph? Do we want to train a bespoke model?” All of these things. And what we decided to do was, “No. Just forget about all this.” These language models are quite good at modeling text and understanding and reasoning about text. So just put everything in a Markdown file or a text file. You’re done. I thought that was quite interesting that we could do that, and there’s a lot of strength in doing that. But also, there are limitations. It’s the way you retrieve and make sense of data that’s extremely large, it takes a lot of work. So I think that technology is getting better. I also do, however, think, there are certain things you just cannot shape just by prompting the model. So to some degree, you do need to touch the parameters of the model itself. So there is this work that I do think does need to happen, and it is happening. The question is, how far can we take it? How do we source data, and how do you also create an ecosystem where people are continuously feeding data to this model so it’s learning about you? Vibhu [00:09:50]: What’s the intuition between why you need to do it in the model? Social Physics and Behavior Foundation Models Joon [00:09:53]: My intuition behind the actual when do you train or even post-train a model versus just prompt a model is if the model has to learn the underlying physics of the world that it’s operating in. So it has to learn new social physics. The places where it doesn’t have to train are the places where it already has the physics. We trust the physics. It already has the base statistics, but it’s just trying to react to an environment. Then I think you can just prompt your way into getting the actions out of it. I don’t think the models that are out in the open have yet learned the complete mapping of social physics of humanity. This is one of the core theses of Simile, right? And one of the core reasons why that is the case is if you look at the data that the model was trained on, these models were trained on the web data, like, whatever was available on the web. And these are really interesting data sets, but they are fundamentally the self-exposed attitudinal data with some behavior data that’s sprinkled around here and there. And it has yet to learn the really deep behavioral nature of people, not just what people say they do online, but what they do in real life. And this is one of what I would consider to be the dark knowledge of humanity that we haven’t quite captured. And it’s these data that would also need to get factored into the model creation. Vibhu [00:11:21]: You call it behavior foundation model. Vibhu [00:11:23]: There’s a good one-liner here, but outside of that, what type of data do you need? What are you changing on the model level? How do you go about modeling, doing a behavior foundation model? The Three Data Buckets: Interviews, Behavior, and Causality Joon [00:11:35]: We think about data in three buckets. So one bucket is interview data. It’s quite interesting. Rich qualitative data is interesting. It’s not behavioral, but we would literally ask people, “Hey, tell me the story of your life.” Vibhu [00:11:53]: It’s just what we’re doing here exactly. Joon [00:11:54]: The question that you all asked at the beginning of this interview literally is the question we also ask. And we ask our participants to go a little bit deeper, than how far I went. Maybe I can give more of my life story in lieu of this. But the reason why that data is interesting is by learning about this very long-tail information about people, you get a lot of texture around this model, like, this person as a model. So even understanding their childhood memory or even their trauma, their first love, these things, quite informative in ways that’s really hard to predict. So that’s one. Then there are two tranches of what I would consider to be the behavioral data. One kind of behavioral data is observational. So these might be like transaction data, or these might be data that you can get by scraping the web, right? So you can imagine why these data sets would be interesting, right, because they give you the base statistics of people’s behavior. Joon [00:12:55]: But then there is the last category of data, that I personally think is perhaps the most important, which is the data that describes the causal mechanism, the whys of people. Some of this is covered by the interview data, the qualitative, because people talk about why they made certain decisions. But really, where you get to see the most behavioral aspect of this is in randomized controlled trials, like RCTs. Imagine you have the same setup, but you have a few different variables that you are trying to tweak. Can you get realistic human behavior out of it in ways where, imagine you had this particular option. Imagine you’re even trying to choose whether you’re going to drink coffee or not. The day you drink coffee versus the day you didn’t drink coffee, does your behavior change? That’s a data set that describes a causal mechanism. This is quite important in modeling people. The reason why this is important is oftentimes when people come to us, or not just to us, but the reason why people are interested in simulation isn’t because they want to predict the future. If you’re trying to win against the stock market, predicting the future is interesting. Prediction vs. Simulation: Shaping the Future Joon [00:14:08]: But most people, most decision-makers, what they want to know is, how can we shape the future? It doesn’t really help you to hear that your sales are going to tank in two quarters. They’re just gonna say, “Wow, that sucks.” What they want to know is, well, what do we need to do now to avoid that future? That’s the causal mechanism. And this is also very hard data to come by, right, because the world is our ground truth, but it happens once. So in a very controlled setup where everything is equal except for one variable, this kind of data set rarely happens. So this is a reason why this data set is both hard to come by and quite important if you’re trying to model human behavior. Swyx [00:14:50]: So behavior, I think, is the hardest data set to acquire. What is out there? What is even possible? You’re not going to know a lot of details about my life. I don’t even have data for myself on my own health or habits, and I just don’t log everything. So how can you have that data? Joon [00:15:14]: So we run a lot of randomized controlled trials. Swyx [00:15:17]: But you put people in the lab, they watch them sleep, or what? Joon [00:15:20]: We do care a lot about the consent process. People know that we invite them to be a member of this community to both share data and have themselves represented in different forms. But we bring a lot of people to the lab, or virtual lab, where we design experiments that would pose them real behavioral decisions. And often in these experimental setups, what makes the difference between what is attitudinal versus behavioral is whether the stake in your decision is real. That’s ultimately what makes it behavioral. So in these setups, we are inspired by our colleagues in social sciences, psychology, and so forth. So when they run studies, the techniques they utilize is imagine there’s an online store that you’re inviting people to come by. Then whatever they purchase in this experiment, they actually get that item delivered. Like, these are the things that make the stakes real. So we run a lot of these experiments, and we also do partner with firms. Right now, we also have customers who are quite excited to at least give us a glimpse of the behaviors that their users exhibit so that we can get a little bit deeper understanding of how people behave in these different platforms. How Customers Use Simile: Populations, Queries, and Experiments Vibhu [00:16:39]: I think on the customer side, they have a lot of data about their users, who has bought. They have the action data. Vibhu [00:16:47]: Can you walk us through an example of what someone comes to you for? What questions would they want solved? Do you customize a model for them? Do you have something off the shelf? What does that look like? Joon [00:16:59]: Today, when people leverage our models, it’s often to better understand the population of their interest. So usually, the start of the relationship, we come together and hear about what population they want us to model, right? So it might be that if you’re a CPG company that’s selling to all of the US, then maybe it’s fairly straightforward. You want to model the gen pop of the US. But at the same time, if there is a vertical or if there’s a market that they’re trying to go into, imagine, they want to better understand, let’s say, people in their 20s and 30s living in California. That’s a much more specific population. So we hear about this population, and we go recruit these people, with consent, and with incentives, and we collect some of their data and create a model of these people. Then what our product allows you to do is query them. So it can take as input a filter that is a description of the population that you want to talk to, just like the one I just mentioned, and an environment. The environment can literally be survey questions, behavioral experiments, It can be A/B testing. Oftentimes, the core use cases are things like concept testing, to start with. But also, people sometimes want to do focus groups or one of the fun use cases that we also serve is even modeling things like earnings calls for public companies. Joon [00:18:21]: So these are the use cases that we often start with. Swyx [00:18:23]: Concept testing, is that an established term? I’ve never heard of concept testing. Concept Testing, Gallup, and Politics Joon [00:18:27]: Yeah. So it has to do with they have, let’s say, different messaging, different products, different ideas. Swyx [00:18:32]: It’s like a marketing exercise. Swyx [00:18:33]: Okay, got it. Got it. Politics? Joon [00:18:36]: We do, have a strategic partnership with Gallup, and of course, Gallup is deep into policy space and so forth. Right now, we have not worked deeply with politics, like that area just yet, however. Swyx [00:18:49]: I’m curious if there is demand or if they really would have different needs that somehow fundamentally don’t mix with your existing, users or people. Joon [00:19:00]: I think there’s certainly demand. Joon [00:19:02]: But we are very much mindful of how this technology gets adopted and the societal impact that we’ll end up having with this technology. And I do see politics as an area where a company has to be particularly thoughtful about the way they operate and make impact. So this is where we also want to make sure that we form enough of guardrail and perspective on how to leverage this technology before we go on to serve markets like the politics. Swyx [00:19:29]: I’ll give people an example. one of my favorite shows is The West Wing. I don’t know if people have watched. Swyx [00:19:34]: One of the key storylines is, like, the president has, multiple sclerosis, but they haven’t. they need to figure out how to disclose it. So they run a poll with a fake governor and ask people to respond on the poll, Counterfactuals, Polling, and When Simulation Is Useful Swyx [00:19:47]: They try to make decisions based on the results of that poll on, like, how well they’ll be received, like where, how should we play this? Swyx [00:19:54]: And I’m like, well, I think those counterfactual things, I would use a simulation for this if I could trust it. Joon [00:20:01]: For sure. Joon [00:20:02]: In that show, how’d it go? Swyx [00:20:04]: In that show, it was, like a foregone conclusion. They were like, “We know it’s bad. We just don’t know how bad.” And then the poll came back. It was like, “It’s really bad.” And then they just did it anyway. Joon [00:20:14]: Part of it is to show, right? So you’re, you’re looking at the idea Swyx [00:20:17]: Maximizing drama. Joon [00:20:18]: How bad could it be? Oh, it’s horrible. Swyx [00:20:20]: And to some extent, I think that is part of the trick of the, or the challenge or with being a customer of yours, which is that if I know it’s. if I roughly know and can intuit Swyx [00:20:35]: What the effect is going to be, do I need you? What sensitivity of it, of effect do I need in order to make a decision, right? So for example, if I, my approval rating is 50% Swyx [00:20:48]: And I, they have this negative piece, news item comes out, and it drops to 30. Swyx [00:20:52]: If it drops to 20, if it drops to 40, do I care? No. It, I know it drops. It’s negative. So when do I care about simulations? Joon [00:21:01]: You do something that’s clearly bad, that’s not popular, and people don’t like you, like, yeah, it’s like Swyx [00:21:05]: You don’t need a simulation. Joon [00:21:07]: Yeah. Well, so there are a couple of things. one is, there are use cases where, like every day, developers, designers, policymakers, marketers, every single day, they create assets. They create new products. And turns out, it’s many of the decisions in hindsight is obvious. Yes, of course this is bad, but we still run those studies because understanding the magnitude and understanding how acute something is quite difficult, even if, we feel like, of course, like this makes sense. this is the reason why we make so many mistakes. Like, every time somebody goes online and say something that has huge backlash, you look at that and like, “What an idiot.” However, it’s tough. That’s one. There’s also another aspect here, which is, again, this is the reason why simulation is different from prediction. In simulation, in the ideal case scenario. So what simulation is trying to show is it’s trying to show each step of the way or each step that we need to take to get to a certain outcome, right? So in the most advanced simulations, sometimes the next step that we’re suggesting might be quite counterintuitive. The analogy that I sometimes give, and I ground it in a more realistic example, but, I, as I mentioned, I’m a huge fan of science fiction, and I don’t know how, many of the audience members have read, like, things like the Foundation series by Asimov. Simulation as a Path, Not Just a Prediction Swyx [00:22:37]: Oh, yeah. We’ve mentioned psychohistory a number of times. Joon [00:22:39]: Okay, fantastic. So I might be, talking to the right crew. If you read Foundation series, literally the first act is there’s a group of scientists who have found out that, “Oh, our galactic empire is going to collapse, and we’re going to have 30,000 years of unrest.” And they run psychohistory, the simulator that tries to teach them, “Okay, how can we keep this unrest to a 1,000 years?” And they plan this out, and the first step of that plan is to get the scientists who say, “Okay, this is coming,” exiled into this random place in this, galax- galaxy. Swyx [00:23:18]: Terminus. Joon [00:23:19]: Exactly. And that’s so counterintuitive. Like, what a strange move that you literally sent the group of scientists who was raising voice around this potential collapse of galactic empire into nowhere. How is that the right first move? Well, it turns out in this particular simulation, that was the move. Joon [00:23:40]: It’s these things, right? And the reason why these reasoning is possible is because you’re showing the step function or each step that results in a particular outcome. So really what simulation allows you to do in its highest form is you give it not a problem or question, like what would people answer to the survey? That’s not what we do. What we tell it is, “Here is a goal that we have. In the context of foundation, we want to keep the unrest to a 1,000 years. What is the path that we need to take now to get to that particular future?” And that’s what simulation allows you to do. Now, translating that into real market, imagine you’re a automobile company and you’re about to release a, EV, and you’re trying to understand, well, how do we market EV, to make sure that our stock price goes up? But what if the answer comes down that, well, you can market your EV in XYZ way, but that might change people’s perception around the cars that’s not EV and make your overall sales to go down. Not very intuitive, especially all you’re trying to optimize is EV salesss, and that’s the only thing that you’re tracking, then that might result in a completely wrong solution, or at least different solution than what you would have expected, whether it’s right or wrong. Joon [00:24:57]: That’s the power of simulation. Swyx [00:24:58]: For listeners, we covered a similar topic with Mikhail Parakhin from Shopify, where they are working on SimGym. I don’t know if he ever talked to you about it. it’s very similar. Joon [00:25:07]: I Swyx [00:25:07]: The goal is increased conversion, but then the journey is very unusual. Joon [00:25:12]: Journey is unusual. Swyx [00:25:12]: Yeah. The-- He’s trying to look for interventions on a shopping trajectory, which is similar to what you’re saying. Like, it’s not about the attitudinal, is your word for it. Swyx [00:25:24]: It’s about behavior. Joon [00:25:25]: It’s about behavior. Swyx [00:25:25]: And that’s exactly the difference, right? It’s, like, not about the near-term direction about-- but it’s more about, like, how do you affect multiple turns of interactions. Vibhu [00:25:35]: You had a good quote at the start about this as well. It’s not about people wanting to know the outcome. It’s about how they can change it, change the way to get there, something like that. But I wanna take it back to how do we know this is grounded? Like Grounding and Evaluating Digital Twins Vibhu [00:25:47]: How do you run evals? How do you test that simulations come through? if I was to do the same thing that you described with, say, your favorite LLM, Opus, GPT-5.6, have some agent to map out these things Vibhu [00:26:02]: How different are the answers we would get if I give it the same goal, the same objective, make a decent system? You’re saying that you need to change the model weight. You have your own solution to this. But how far off are we, and how do you check if it’s grounded? you have some interesting stuff on your site that points to how you run real evals, but if you could take us through that side. I think that’s one of the big concerns that people have. They’re like, “LLMs hallucinate.” Vibhu [00:26:27]: “You’re just hallucinating layer after layer,” right? Joon [00:26:30]: The way we do this, and this is the paper that we worked on after the generative agents paper that really became the, at least for Simile and also the field of simulation and synthetic panels, really became the foundation. Yeah, this is the paper. the paper is called Generative Agent Simulations of 1000 People. Here’s what we’ve done. For this paper, we brought 1,000 people that’s representatively sampled from the US to a virtual lab. And what we have done was we spent two hours collecting fairly wide-ranging data. In this particular study, we focused a lot on this interview data, that was, whose script was taken from this project called American Voices Project. And then we would also pair that with a lot of behavior data and so forth, whatever we can collect within two hours. And then we would send these people away for a couple of weeks. And during that time, I would use this data to create their digital twins. And I would bring the humans, participants back after 2 weeks and have them complete a battery of surveys, experiments, behavior studies. So we have the list here, which included things like behavioral economics games. We would run literally, like, Big Five personality test, General Social Survey. We would also go ahead and run the randomized controlled trials that were published on PNAS. And we would have their digital twins predict how the source individuals would have acted in these studies and surveys. And this is where we could replicate people’s behaviors and attitudes 85 percent as accurately as people would replicate their own. So that was the first really paper that gave this validated results that we can model individuals in an accurate way. And what we ended up finding now, of course, in AI space, so this paper came out at the end of 2024. AI space, a year and a half, 2 years, that’s a lifetime. 85% Accuracy and Why Frontier Models Miss Human Behavior Swyx [00:28:24]: Yeah. Just, for listeners who are not seeing the YouTube, I just wanna say, like, the headline figure is 85 percent accuracy, like, which is a big improvement over all the other Swyx [00:28:34]: Methods that you showed. Joon [00:28:36]: But the part that was particularly striking to us, especially as we improved this technology even further, was the generative AI models like ChatGPT, Claude that’s coming out, it does give you the right foundation. However, what they do not consider is the true attitudinal and behavioral aspect of people, especially in the population that you care about. So what these models are really good at today is they’re trying to become the super rational, objective machines, right? So you go get their data from places like Mercor, Scale. You talk to professional programmers, scientists to create model that’s amazing at reasoning. That’s what they do. Simile doesn’t care about any of this. The models that we’re talking about here, what we’re trying to create are models that are as dumb as I am, right? So if I make some mistakes, the model has to make the same mistake. Swyx [00:29:34]: Oh, that’s very hard. Joon [00:29:35]: That’s very hard. Swyx [00:29:36]: You’re solving Murphy’s paradox. Joon [00:29:37]: That’s exactly. And this is a completely different data and training objective. This is also where we see quite a bit of discrepancy in the performance in human behavior prediction between the frontier models, Simile’s model, and the models being created in this space, where in some cases, the model performance of frontier models go all the way down to 20, 30 percent, especially if you go into that more niche population on topics that our customers would care about. On more gen pop, it might be around 50 to 60 percent. So it’s not very robust. Like, you wouldn’t want to make your decision off of these and these findings. If you can bring that up to 85 percent, that is ultimately what people end up getting very excited about. Swyx [00:30:20]: Yeah. Do we wanna keep going on the paper, routes? Joon [00:30:23]: Yeah, for sure. So the last one, was an interesting one. So this, paper was the follow-up paper that we had, to the 1000 agents paper, where the idea was now can we augment the models even further and post-train a model based on a lot of randomized controlled trials? So this was an interesting one. The data is always the most interesting part of modeling in many ways. The data that we got here was there’s this, there’s this platform called Open Science Framework. So some, the audience might be familiar with this. And there has been, especially in the social sciences over the past 5 years or so, there has been this concern around replicability of studies. And so it was a bit of a crisis, the scientists acknowledged, where we rerun the study and we don’t see the same finding. Post-Training on RCTs and Replication Studies Vibhu [00:31:12]: Oof. Joon [00:31:12]: It’s tough. And the reason why it’s there-- that was often the case was there’s this survival bias where the papers that get published often need to maintain what we call the value of less than 0.05 in the experiments that we ran. That suggests that only-- there’s only 5% chance that the results that we saw is false positive. But the tricky part was all the papers that were not published, and there’s still a 5% chance that whatever we publish is totally just randomly generated. Like, there’s a 5% chance that, hey, this effect is not real, but it just happened to be real because of the sampling bias. So because of that, what scientists started to do was they started to register their studies. So before running an experiment, they would go to this platform and say, “Here is the data. Here is the population that we’re collecting, and here’s the hypotheses.” And they would just say, “Here is our hypothesis.” Like, “This is what we believe.” And you cannot retroactively change those hypotheses. This is what gives us more scientific statistical confidence that whatever effect that you ended up seeing is true. So that ended up creating this really interesting platform where there’s one platform that has now contains tens of thousands of real-world experiments and hypotheses. And a lot of these are really high-quality, like, professionally designed behavior studies and random- randomized controlled trials. So we got the data and the studies from this platform and used that to make a point. And this particular, model is not, something that we’re serving commercially because this was a part of the open science. But this particular data set, helped us make a point that by collecting a lot of these randomized controlled trials, that are really well-designed, we can make significant improvement in model’s capability to predict human behaviors. So that’s what this paper was about. Vibhu [00:33:10]: Is this stuff done on a individual level? Like, do I need to tune the model per individual, per company? Is there foundation model changes and then some slight post-training? Anything you can share there? Population-Level vs. Individual-Level Models Joon [00:33:21]: So this particular model was trained. the data we had at the level of individuals, but this particular model was trained. We experimented with both. And this is what we end up doing at Simile too. We always train 2, distinct model. One is what we call the population-level model. The other is what we call the individual-level model. And both take very similar input, which is the description of a subpopulation or individual and a stimuli. In this particular work, we’ve done the same. Here, the results that we are reporting are much more geared towards individuals because we do think that is a harder task in many ways, but that’s what we have done. Vibhu [00:34:02]: You seen anything on the questions that humans can solve that models can’t solve? So like Human Biases, Mundane Choices, and What Models Miss Vibhu [00:34:09]: Currently, it’s, I live 5 minutes walk away from a car wash. It’s a 10-minute drive. Should I walk or drive? Joon [00:34:16]: Huh. Vibhu [00:34:16]: The model will say, “Oh, walk to the car wash.” And, you don’t have your car. Vibhu [00:34:20]: Is anything like this a problem in simulation? You would assume, like, very simple for human to think about, but if the model is saying you should walk to the car wash, anything here? Joon [00:34:32]: It’s less, what can we solve, but I think it’s more about what biases or mistakes do people make that models miss. Like, imagine that you are, like the. When I was still at Stanford, I lived in Palo Alto. So it’s about, I would say, 40-minute walk from the campus. You ask the model, “Okay, let’s go home. What can I, what can I do?” It would likely call an Uber or, give me, the bus time. But for the longest time, I really liked walking back. And the reason why I wanted to do that was not for efficiency. It really helped me think. And I like to walk for, half an hour or 40 minutes or so a day, where I just get to, just think about ideas, research, just get lost in my thoughts. That’s very human activity. Unless the model has seen that and understands the importance of that activity, it would miss these kinds of features. So that I think, is fundamentally what we’re trying to model. Like, what is fundamentally human might not be the most efficient thing to do, might not be the right thing to do, but things that make us who we are. Swyx [00:35:43]: I’m curious if, there are some data sets that you really want that would materially help you. One version of this may be interesting, which is more valuable to you to acquire as a data set, all of LinkedIn, all of Twitter, all of Facebook? What Data Matters: Social Media, Transactions, and Facebook Joon [00:35:57]: It’s a little bit hard to rank, in part because, there’s, there’s this product saying where no feedback is wrong because it teaches you something about your users. Doesn’t matter what feedback. Joon [00:36:11]: I think it’s a little bit like that. Swyx [00:36:12]: So just whatever is bigger. Vibhu [00:36:13]: What about a different domain? Say it was. What about all of Amazon data? Joon [00:36:17]: Oh, yeah. Vibhu [00:36:18]: Shopping data, right? Joon [00:36:18]: Shopping data. So Amazon data is interesting in that it’s very much behavioral, although, like, what people do on social media, you could squint and say that is also behavioral. But the transaction data is always interesting. It is also most commonly available, however. Joon [00:36:33]: If we were to look at purely social media, like if you really, if I were, if I had to really pick, Facebook likely is interesting because I do think it is most a default version of people. Because you go to LinkedIn, it’s very much professional environment. So people put up their, they have their guards up, right? And that still is interesting because that is true human attitude and behavior, but it is not your base state. you go to Twitter- Twitter, people have their own crazy personas, or depending on who you are. Like, my Twitter profile and, persona is very much, initially was I was very much an academic. “Hey, I’m here to share my studies.” Now, I share, things that’s related to Simile. But Facebook is one of those more private space where people just connect with their friends. In that way, I do think it shows you a little bit more about who that person is. So if I had to pick, I’d likely pick, Facebook. Swyx [00:37:30]: Yeah. And you’re interested in, like, the whole person and their background and philosophy. I, is it too clinical or too machine learning-oriented to just say this is just ways to inject variance and biases? The broad question, is, like, is this any better than a randomized, like, combinatorial explosion version? So we have a link to the Tencent Billion Personas, Synthetic Demographics, and Bespoke Data Swyx [00:37:54]: Billion persona paper, where they did not do any of the groundwork that you are doing. Swyx [00:37:59]: They just did like a cross matrix of here’s all the professions in the world, here’s all the people, possible backgrounds in the world, do a dot product across all of them, and that’s it. That’s your prompt for a billion people. Swyx [00:38:12]: This will do something. I don’t know if it’ll do what you do, but it gets you some way, some percent of the way there. Joon [00:38:18]: So this was an interesting paper. Like, what I admired about this paper when it came out was the scale. And you do gradually want to be able to simulate really large societies and interactions. So the scale is definitely admirable. it is relying heavily on the known statistics that went into training the model. So to the extent that you believe that statistics is correct, this is not a bad way to go about this. But the thesis here, and this is something that we also have seen in the market, like if this works, then we have solved simulation. Joon [00:38:54]: It, Swyx [00:38:55]: Because I survey, like, okay, 5% of the US population is in construction. Swyx [00:39:01]: The other 5% is in medicine, whatever, right? And then you just keep going down the list, and then you do the other side. 5% has, like, the big 5 personality Swyx [00:39:08]: Of, like, neurotic or whatever. That’s it. Joon [00:39:11]: That’s it. So if you believe that the underlying data set and the platform that we’re leveraging has all the right statistics, then this will have solved it. you’re at that point merely retrieving the knowledge that is already embedded in the model, in the model parameters. That’s not, unfortunately, what we see, where there is such detailed and also niche knowledge about people that if you just take one example, it might feel very mundane, but it’s quite rich when you put together, that you do need to do a lot of bespoke data collection to better understand people. And this is also, I think what makes this particular, job fun, which you want to deeply understand people, and the process of deeply understanding them requires a lot of attention to the details. And you do need to pay attention to and pay respect to the daily lives that people lead. Scaling Simulation: From Thousands to Societies Vibhu [00:40:04]: I wanna talk about scaling simulation. Vibhu [00:40:07]: So what can’t we simulate, what can we simulate, and how does scaling affect this? So how big are the models? What if we go from, 8B, like, couple 100 billion Vibhu [00:40:18]: Like billion000 parameters, billion000? Do we get scaling? Any interesting emergence? Like, at a certain scale, at a certain amount of training, you uncover anything unusual and any learnings from that? Joon [00:40:31]: What we are seeing is at Simile, so we do post-train our own model. The thing that we’re seeing is the early glimpse of scaling law in simulations. The more data about humans and more compute you ingest, you start to get predictive and predictable gains of the model performance in simulating it, simulating people. Vibhu [00:40:51]: Ooh. We need a scaling law curve. Joon [00:40:52]: It’s scaling law. Whenever you find it’s a beautiful thing. And we’re starting to see the glimpse of it, which is quite exciting. But if you talk about the ambition of simulation as a whole, it’s not merely about building a model. It’s about building a model, then creating the agents that become the individuals in a much larger ecosystem. So they’re creating this multi-agent simulation. Down the line, you want these multi-agent simulation to also live in a very rich environment, right? What we are really trying to get to at that point is, hey, can we create. All right, let’s do a time machine game again, and 5 years, 10 years into the future, can we create a simulation of 8 billion people living on Earth? I think that’s quite interesting. And that really is the vision. And once you get to that state, the questions that you can help answer for the society also start to change from my perspective. The answers are fundamentally about emergence of the emergent behavior of society and large groups of people. Joon [00:41:53]: So the questions that I get excited by, and maybe this is a stodgy- a bit. I have my, academic side of me. Joon [00:42:01]: And for me, it’s questions like, can we help solve climate change? If you look at climate change as a problem space, this is what we, like social scientists would often call it the wicked problems, problem where you have many actors with competing incentives for trying to make a very complex decision and coordinating that coordination decision. Very difficult to really solve in real life, which is also the reason why we couldn’t solve it. Can simulation help us solve that? Another one is, can we understand the signals for collapsing democracy, or can we understand or can we uncover the origin story of the monetary system? These are societal questions that we never really had a good way of answering. If we can create simulations of our society, you have to believe that these are the problems that we can solve. So that’s really the ambition of this field. And, I also think, yes, I think there’s a Nobel Prize to be won there, which wouldn’t be surprising. And I think there’s some amazing societal impact that we can have to help people make better decisions. Climate Change, Democracy, and Societal Simulation Swyx [00:43:04]: Nobel Prize in economics? Joon [00:43:06]: In economics. Swyx [00:43:06]: Oh, I see. I see. Rooting for you to write that paper. Joon [00:43:10]: One of these days. But, one of the scholars that I was deeply inspired by, When I was coming into the space of simulation, is this scholar, named Thomas Schelling. Schelling, Agent-Based Models, and the Nobel Prize Swyx [00:43:23]: Schelling point? Joon [00:43:24]: So the canonical example of the work that he’s done was he was one of the creators of agent-based modeling. So this was, like, in the 1970s and 80s. It’s very early days, but this was truly one of the first exemplars of simulations. And one of the canonical model from that time, and of course many of these simulations are trying to tackle the societal problems that’s most relevant for their era, it was called the model of segregation. So racial segregation was a big topic, that, we cared about. And what they’ve done was they created this grid world where they had red dots and blue dots. And these dots were, back in the day, like, they were the agents, and they had a simple rule that governed their behavior. If certain percentage of your neighbors are of different color and if that goes above certain threshold, then you move to a new location at random. Joon [00:44:21]: One of the striking finding of this paper or this agent-based model was for the longest time, people thought the segregation within society was caused by explicit and overt racism. Joon [00:44:34]: But if you look at this model, people’s preference towards living with people of the same color, that preference can be very minute. Joon [00:44:42]: But the very small difference causes the society to segregate completely over time. This was very counterintuitive for a lot of people. And this particular work ended up informing housing policies. Mixed income housing, got really inspired by this work. And Thomas Schelling ends up winning the Nobel Prize for having laid the groundwork for very early versions of simulations. The opportunity that I do see here in the more scientific terms, is agent-based models for the longest, had impact in the 1980s, 90s, to some extent, early 2000s, but it has now gotten forgotten by the community a little bit. Because as you can imagine, red dots and blue dots is not really a rich description of people. Joon [00:45:31]: But with the emergence of things like generative AI and, in particular, generative agents, we do have an opportunity to create these agent-based models that are high fidelity enough to help us make really complex decisions. And that’s the opportunity that I see. If that truly works, then yes, that is the work that will result in a Nobel Prize. Swyx [00:45:53]: Yeah. For what it’s worth, and I grew up in Singapore. 80% of Singapore is in public housing, and public housing has, enforced racial quotas for exactly that reason, which is very interesting. okay, so we talk about scaling, we talk about all these, the agent possible applications. Cost, Reuse, and the Economics of Simulation Swyx [00:46:13]: I’m scared about the cost. if you even-- let’s just keep it to the US, about 8 billion people. Swyx [00:46:21]: But, how much does it cost to model so many hundreds of millions of people? Joon [00:46:26]: Oftentimes today, we don’t start at that scale, this stage of the, of industry and simulation as technology. But we can get our users extremely rich and meaningful insights even by modeling thousands, tens of thousands of people. And today what we do is every week we are collecting data on the scale of tens of thousands people’s data, and we have panel partnerships that gets us to tens of millions of people globally. So that’s what we do today. Swyx [00:46:55]: And just as a side note once you’ve collected one person for one study Swyx [00:46:59]: Can you reuse that same person for all the subsequent studies? Joon [00:47:03]: That’s exactly right. Swyx [00:47:03]: Okay. Joon [00:47:04]: The beauty of this model and these agents is the fact that they are domain-agnostic. Joon [00:47:08]: That what you’re really trying to understand is what is the fundamental nature of these people? What’s their social physics? And there are a lot of, a lot of, people that does change over time. Like, even, like, even things like, how many times have you gone have you been to, like, CVS the past week? that will change. But there’s so many traits about people that are also known to never change. Like, your risk tolerance doesn’t really change over time. It’s very consistent. So it’s these things that we’re trying to learn. But the scale we are operating is right now hundreds or, tens of thousands to hundreds of thousands. And in many of the core use cases that we are deployed in, and this is more than enough population, to cover those. Really, at that point, what you care about is less the number of people, but more do you have the right subpopulation of interest covered? And this is also the reason why people want a larger sample. It’s not because they want, stronger statistical guarantees. It’s more that can they filter down to any population of their interest. However, you can also imagine in 10 years, if we truly believe that the compute is going to scale, that we’ll have much more availability for compute, and our ambition for simulation is also going to scale accordingly, there’s definitely a reason for us to create an entire data center worth of simulations. Joon [00:48:35]: Or in my hunch here is I do think in the next some number of years, we will start creating simulations that will cost as much as training a foundation model. But perhaps it’s going to be so valuable to the society that it would be a no-brainer. Right now, even today, like, we are training bunch of new foundation model just so we can say we trained one and we spent tens of millions. But if we can create a simulation at the level of society that would solve climate change, I would run that today. I would raise the money right now just to run that. Multi-Agent Simulation and Social Influence Swyx [00:49:10]: Amazing. the follow-up question is, does it also compound if you let the simulations talk to each other? Swyx [00:49:18]: Or do they already do that today? They don’t, right, as far as I understand? Joon [00:49:22]: It depends on what simulation you’re trying to run. Joon [00:49:24]: In the multi-agent simulation setup, the agents do talk to each other. Swyx [00:49:28]: Right, which is exactly Smallville, right? Joon [00:49:29]: That’s right. Swyx [00:49:30]: But a lot of times, for example, in commerce, you’re just by yourself, so there’s no point talking. which is way cheaper. Vibhu [00:49:37]: But they use all these levels, right? Like, you decide what you will buy based on what other people around you buy and talk about, right? Swyx [00:49:43]: It depends. Vibhu [00:49:44]: It depends. Swyx [00:49:45]: Again, I’m, I’m coming at this from a cost point of view. I’m like, “Oh my God.” Like Vibhu [00:49:48]: I think Swyx [00:49:49]: If there is, like, some combinatorial thing of, like, thousands of people talking to thousands of people, then that one million X’s might cost. Vibhu [00:49:56]: I have a very different view as the cost point aside. Like, running these studies in reality is a lot more expensive, right? Running any study like this is you gotta have people do it, you gotta sign people up. It’s very expensive and sometimes, like, not feasible to run the study. Vibhu [00:50:14]: But the outcome or the decisions you make are very expensive on them, right? So spend X million on something that, the overall process costs 100 million might as well, right? There’s, there’s a lot of value to be had there. It’s a small cost, but I’m excited on the cost side. Joon [00:50:33]: To some extent, and when you deploy technology, you often want to deploy in a way where you can replace existing budget or you can make things more efficient, and that is the best way to deploy. However, the way you capture the long-term value of the technology is making the argument that, no, it’s the upside, that by making this better decision using simulation, you have saved yourself or made yourself hundreds of millions or even billions of dollars, and that’s a case to be made. Vibhu [00:51:06]: Random tangent question. So if you’re doing a lot of inference, a lot of model multi-agent stuff, are you at the point where it makes sense to, train a model that’ very sparse? You’re expecting to do multi-million dollar runs. Are you thinking about this in model architecture standpoint or inference efficiency, or, you’re still at the research phase of it works, we’re not super there yet? Joon [00:51:34]: Efficiency, we do think quite a bit about. this is technology that is deployed now in some of the largest enterprise companies in the world, and we do process significant number of queries, that are trying to, simulate the populations in the world. So efficiency is a consistent thing. we don’t want to over-optimize too early, so I wouldn’t say, like, this is the higher bid Right now, but this is definitely something that we think pretty carefully about. Swyx [00:52:05]: Yeah. Are there other case studies? So we, you talked about CVS, talked about Gallup, Deloitte, Wealthfront. Efficiency, Enterprise Use, and Real-World Case Studies Joon [00:52:12]: Wealthfront is an interesting one, because one of the things they were trying to do, they were one of the first customers that wanted to do product testing that goes beyond just asking people what they think about, let’s say, behavior experiments and so forth. So there, really what we had to do was reason about multimodal input, so images, but also you can also imagine, like, these agents traversing through Figma mockups or websites. So some of the things that our agents can also do is it can be given a domain, like, or, like, a website URL and go use it for a while. It’s these things. And Wealthfront was one of the first, customers, that was very excited about this possibility. Vibhu [00:52:53]: What have people been asking? Like, is there any demand that we have not covered? Like, UI testing, right? Vibhu [00:52:59]: I wanna try a new. I wanna ship a new feature, test the UI, simulate how people will do it. Any interesting things that you’re seeing demand for? Product Testing, Websites, and Synthetic Panels Joon [00:53:08]: Today, a lot of the demand does come from like, the places where people have historically used human panels, we can now replace with agents, and these synthetic populations. And this is not replacing human panel. in many ways, the simulation that Simile is building is grounded. So the way that I think about this is we are trying to represent humanity at scale. And in that way, the use cases are what we would expect, but it’s the scale of deployment that surprises me. Joon [00:53:44]: Turns out there are so many decisions that people make every day in these organizations, groups, and we want to be able to say, “We listen to people. We have consulted our users.” But in reality, that is rarely the case because getting to people and asking them many questions, it’s difficult. It’s both costly, time-consuming, but most importantly, people are just not available. If I had to answer 1000 survey questions for this one particular, vendor, even if I wanted to do that, like, I would never do it. And that’s very much the case. What simulation can do is ensure that the voices of people are always represented in rooms where the decisions for them is made, right? So all the stakeholders of this particular product launch, ideally they’re consulted. That’s what this technology really is trying to enable. Market Size, TAM, and Human Decision-Making Swyx [00:54:39]: In my mind, that means it skews towards more consumer focus, right? Like, anything with a wide enough customer base where you do benefit from the diversity that you represent. What are some rough statistics, just for people who are not familiar with this market in general, what’s the market size that. I’m sure you have some, like, rough numbers. market size is, like, a vague question Swyx [00:55:01]: But, like, how much do people spend? Joon [00:55:03]: So market research is a $100 billion industry. Joon [00:55:06]: But the thing about simulation is not a tool for market research. Simulation is a tool for human decision-making. So the question around what is a TAM here is quite tricky, right? Because it’s easy to say, “Well, market research TAM is roughly 100 million or 100 billion.” so is it a TAM? And not really, right? Because in many ways, you’re trying to inform all human decision-making. You’re trying to inform every decision that are made about humans for humans. What is a TAM for that? It’s really unclear. And I’ll be honest. Like, I have a scientific background, I have a research background, so I didn’t come into the field calculating, oh, what is the TAM for human decision-making? But I just had to assume, well, if we can inform every decision that is made about human for human, that has to be big. Swyx [00:55:58]: Some- something valuable. Joon [00:55:59]: Exactly. Swyx [00:55:59]: To some extent, you are a unicorn founder now, and you have to care as a CEO. But, like, I do think, like, yeah, when you go into these boardrooms with people that you’re quoting millions of dollars of contracts for, like, you have to say, “Well, here’s what you spend on humans-” Swyx [00:56:15]: “. And here’s what we save you, and it’s 85% similar.” Joon [00:56:19]: And certainly, the value case, is something that we care deeply about. Like, what is the value that we provide to the users and the decision-makers? But this is also where, like, as a founder, I think valuation only tells one very superficial aspect of the story, and I try not to think too much about valuation, in general, because that’s not what also motivates a team or certainly doesn’t. I’m, I-- Again, the interesting thing about researchers is we are happy living in academia, getting paid next to. we get paid okay. we don’t get paid that much, as a researcher here in academia, but it’s the impact and it’s the, it’s the value that we can provide to the individuals and the society that really drives us. And in that way, ultimately what drives us is the impact. Does the simulation we provide have a real impact in people’s decision-making in ways that progresses our society forward? If the answer is yes, then yes. that has to be great business, and we see that in numbers, and we do care deeply about that upside story, but that’s the heart of it. Where Simulation Goes Next Vibhu [00:57:27]: Do you have any timeline predictions? So we talked about scaling laws of simulations. Vibhu [00:57:33]: You brought up, okay, maybe one day we can simulate how to solve climate change. Vibhu [00:57:38]: Where are we now? Vibhu [00:57:40]: If that’s not the end state, what is an end state, and what does progress look like? Joon [00:57:45]: So what I sometimes tell people is simulation as industry, it feels a lot like where GPT-3.5, GPT-4 was, for the AGI saga, which is we have now technology that is powerful enough to do real damage on the verticals that we are tackling. At the same time, there’s a lot of progress that is yet to come. And that’s, I think, where this is. So the way I see it, I do think there will continue to be breakthroughs both in data, in algorithms, and there will be much more aggressive scaling that will also happen over the next few years. But I think that’s roughly where we are. Swyx [00:58:27]: I think that was about the rough set of topics. Anything else that we should have asked you or you wish people asked you more about Simile? Simulation as Painting and Understanding Human Essence Joon [00:58:38]: I think the, what’s, for me, what’s quite fascinating about simulation, it is very impactful technology, but it is also very interesting technology, both in terms of, like, what it means for human society, our philosophy. And the way I sometimes interpret simulation is. So going back to my background, I as I mentioned earlier, I started my career as a painter. it was a professional pursuit, and I did oil painting, for figures. So I got my training originally in the realism studios, and that’s what I spent a lot of my, years, doing. Simulation is a lot like painting, right? The best paintings teach you something deep about the subject that you’re trying to represent. And it is always not a perfect representation. It-- No painting is perfect. There’s always some small differences and discrepancy, but what it does is it tries to highlight the thing that matters the most about the subject. Swyx [00:59:47]: The essential Joon [00:59:49]: The essential essence. Swyx [00:59:49]: Yes. He, you, he’s brought up some of your work. Vibhu [00:59:53]: Just nice to put it up. Joon [00:59:54]: Yeah. So these are some of the works. So this is from, my, personal website that I maintain when, I was still a researcher. Swyx [01:00:00]: I think a lot of people will say, like a Picasso, like anything postmodern is, like, very much focused on the essence. Swyx [01:00:09]: Right. yeah, but I don’t know if any one of these evokes something that you like to tell the story of. Joon [01:00:15]: No, it’s one of those things where, each of these paintings, drawings, whatever it may be, it is trying to surface something about the subject that you feel deeply about onto the surface. when I was a painter, and artist, the topic that I cared really deeply about was, the more mundane aspect of human lives. This shows up in some of the, some of the work that I’ve done, where, like, I did this entire study of a rural town where I went around and took photos of people for not really doing anything special, but just living their everyday lives. I thought that was the most interesting thing. I’m somebody who has this perspective where, the world is oriented around this fractal shape, and you have two choices to understand the fractal shape. You either go outward and try to explore as much as you can to understand the broader shape of the fractal, or you go inward because, the outward resembles the inward, shapes. And understanding the mundane aspect of it was very much that. Simulation has a lot of this, right? You’re trying to understand even the most mundane aspect of people. When put together- teaches you something really deep about that individual and the society. So I think that’s what’s interesting about simulation, the way, the same way that AGI helped us better understand or really think critically about humanity and human intelligence, simulation is really an exercise of understanding more about human society and our collective lives. So that I find to be, yeah, particularly interesting. Swyx [01:01:56]: Yeah. Now you’re reminding me that some of the best biographers, documentarians, and even photographers, they’re taking a photo of you. Swyx [01:02:05]: But before I take a photo of you, I must spend-- I must, like, follow you for a week just to understand you? Swyx [01:02:11]: Which some artists, some do. Part of your work, there’s a very famous book called Working. I don’t know if you’ve, been referred to it before. Swyx [01:02:18]: It’s very famous, like, to the point of having a Wikipedia page Swyx [01:02:23]: About this like, really depth understanding and interview of people as they, about their lives, which seems mundane, but is told in a very, compelling way. Yeah, 1970s as well. Joon [01:02:34]: Okay. It was an amazing decade. Vibhu [01:02:39]: Before closing question UBI, Future Questions, and the Value of Simulation Swyx [01:02:41]: Okay, here we go Vibhu [01:02:41]: You said that you started Simile with your 10-year question, right? If we do that now, 10 years down, what can we simulate? What would you simulate if, like, if you’ve made significant progress, are there any questions outside of the ones that we brought up? Any- anything that you think is most impactful? Anything that you would go vision 10 years out? Joon [01:03:03]: In many ways, as I mentioned, I am somebody who is very much impact-driven. So the what would inspire me is I would want to ask, 10 years later, what would be the most important societal question that we as a society have to ask? I would love to tackle that. Like, do we need UBI? That could be an interesting one. Swyx [01:03:24]: Ooh, has anyone done that? Joon [01:03:25]: Well, we were thinking about it. Vibhu [01:03:27]: Can we get access? Can we just Swyx [01:03:28]: So OpenAI, this is, like, just trivia now. Like, OpenAI, or I think Sam Altman funded a study on this Swyx [01:03:35]: In Africa, and the answer was no. Joon [01:03:37]: The answer was no. But, what, was it something about the implementation? Swyx [01:03:41]: Yeah, I know. It was a skill issue. Joon [01:03:43]: Or was it something about, But this is the thing. See, when Sam Vibhu [01:03:46]: Funny news article Joon [01:03:46]: Altman funded this particular, Swyx [01:03:50]: He spent 14 million dollars? Oh my God. Vibhu [01:03:52]: It’s a little more. Joon [01:03:52]: Quite a bit. But this is the thing. This is the reason why you want to run a simulation. You spend 5 years, 40 million dollars on this one study and have one finding, but if you can run simulation many times instantly, then that’s the value. Swyx [01:04:07]: I feel like that one could-- you could have done in a simulation. Like, if you can do the housing study, you can do the UBI one. Like, I, come on. Vibhu [01:04:13]: I think sometimes people will spend the money because they wanna verify what you think, right? Like, sometimes you just wanna. Is it right? Like, you gotta test it. Swyx [01:04:23]: Okay, closing question. What are the chances we are in a simulation right now? Are We Already in a Simulation? Joon [01:04:28]: So it’s a fun question, and I assert at some point I just answer, yeah, we’re definitely in a simulation. But what I do, feel, however, is, whether we are in a simulation or not, that, I don’t think that makes our experience any less real. And I think that’s fundamentally, like, what I believe in. Maybe we live in a simulation, maybe not, but for Swyx [01:04:48]: It’s real to us. Yeah. Joon [01:04:49]: Yeah. For me, I don’t really care. Swyx [01:04:50]: Yeah. Unless you die and you wake up in, like, the level higher or below. Joon [01:04:55]: That would be interesting. Vibhu [01:04:55]: I feel like you wouldn’t care. Once you die, then you find out you’re in a higher level. Joon [01:05:01]: I worry about it when I die. Swyx [01:05:04]: I think the other thing that. Okay, so I like the mathematical answer to this, which is, like, the, sheer number of possibilities that you are in a simulation far outweigh the sheer number of possibilities that you’re not. Swyx [01:05:16]: Except for the simplest answer, which is, it is computationally very expensive to have you be a simulation. okay, great. You’ve been very generous with your time. Congrats on all your success. I met you just after your Smallville paper and had no idea that you could build, like, such an enormous company. And then now you’re like, “Well, it’s a $100 billion market, but that’s just where we’re starting.” So this is, very exciting. Vibhu [01:05:42]: I think $100 billion market was not the term. That was only part of it. Swyx [01:05:45]: Yeah, exactly. It’s, if you’re thinking too small. Joon [01:05:48]: Well, I do believe that, maybe my final note here might be, again, I love science fiction. You look at any advanced civilization in science fictions, there’s 2 twin pillar, technology. One’s AGI in some form, and the other is simulation. So I think the market’s pretty big here. Simile as Research Lab and Product Company Vibhu [01:06:08]: Tell us about the company. You guys just raised a lot. You’re half a research lab, half a company. you’re hiring. Where are you based? Joon [01:06:15]: Yeah. So we’re based in Mission Rock, so not too far away from, where we are right now. So we’re in SF, but we are also bicoastal. So we have our, team. I would say our headquarter is in SF, and we have a lot of our technical talent in SF, and we do have a smaller office that just opened up in New York. We are, as a company, an interesting one in that today, there are AI neo labs and then there are AI product companies. Simile truly is both. So this is a company that was founded by 4 founders, myself, Michael Bernstein, Percy Liang, Lainie Yallen. Michael, Percy, and I are all researchers. So of course, Michael was one of the authors of the ImageNet, kickstarted the AI revolution back in 2013, has been instrumental in human-centered AI. Percy coined the term foundation model, and is a, one of the greats of the AI researchers today. And Lanie is my business counterpart, where she led some of the fastest-growing AI native companies from their seed to A and B. But we have this DNA at the company where the vision of the technology that we’re creating is continuously developing, that we are getting people who were my lab mates. We are about 60 people right now. Joon [01:07:28]: 15%, almost 20% of the company population are just my lab mates from Microsoft Research lab. Joon [01:07:36]: And we It’s quite fun because many of them then had gone on to OpenAI, Google Gemini, and these places. And so it’s been a few years since we really got together and had a chance to work together. But now they’re coming back and really building out this vision that I find to be quite exciting, and that excitement is shared. So there’s that motion at Simile where we are a group of researchers trying to do something that no one is working on that we find to be the most impactful potentially. But at the same time, this is, again, technology that can make impact today. So we have an amazing group of engineers, product people, and designers, who are sitting here with us trying to imagine what does it look like to help people understand what simulation can do and make real-world decisions with this. Having both and then deploying it to some of the largest customers in the world today, it feels quite unique. Swyx [01:08:30]: Yeah, it’s very compelling. One part of it was this is the call to action. Like, who are you hiring? You’ve done part of it, which is you have-- you’ve got a very talented group. Who are you hiring? Like, what roles? Hiring and Closing Joon [01:08:41]: So honestly, at this point, we’re hiring across Swyx [01:08:43]: Everything Joon [01:08:43]: All, section. we are always excited to bring on, amazing research talent. Joon [01:08:49]: So if you’re interested in working with, our lab mates, we are always welcoming of amazing, researchers. But also we, hire, amazing engineers, that some of whom I, like, I respect the most. Many of them come from places where we have personal connections with, so many of the members are from Figma, Notion, Rive, and so forth, but also more broadly from the companies that we as a team have really admired. So engineers both in the product side, infra side, we’re all looking for those hires. Swyx [01:09:24]: Well, lots of people. I think you made a really good case. So thanks, and, we’ll see you in the simulation. Joon [01:09:30]: Amazing. Joon [01:09:31]: See you all there. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe

There are roughly 100x more people who use code than who can write code. As code that “just works” becomes easier to generate, this group may be the biggest prize of all — if you can get the agentic interface right. A key trend we have been tracking over at AINews is the absolute explosion in Codex usage this year, with MAU now up >10x from Jan 2026. Less than two weeks after their July 9th launch, OpenAI said ChatGPT Work and Codex had reached 10M users combined (as we cover in the pod, Codex now powers ChatGPT Work, so all ChatGPT Work users are now users of the Codex harness, even if they aren’t traditional engineers) — showing the early innings of what happens when you graduate from coding agents to knowledge work agents: We’ve been calling out how coding agents are “breaking containment” to do everything else this year to power every other part of knowledge work - and it started with the org chart, with a major reorg last month that amounted to two of Codex’s most prominent leaders, Greg and Tibo, taking responsibility over product and ChatGPT specifically, completing a “Superapp” consolidation cycle first discussed in March. With these updates Codex is no longer just a coding tool. In June, OpenAI said knowledge workers already accounting for roughly 20% of Codex’s user base and growing more than 3x as quickly as developers. A product dedicated for knowledge workers was being pulled out of the Codex team. However, knowledge work has a different set of problems and environments than coding. For decades, knowledge work has been scattered across different primitives like documents for writing, spreadsheets for analysis, slide decks for communication, and specialized applications for everything else. ChatGPT Work now enables users to work across every primitive with agents. Instead of opening an application and manually operating its features, the user can describe an outcome and collaborates with an agent that can assemble the tools, context, and artifact needed to reach it. From building no-code products at Airtable to leading Productivity Engineering at OpenAI, Akshay Nathan has spent much of his career trying to make the power of software accessible to people who do not write code. In this episode, Akshay joins swyx and Vibhu to unpack the launch of ChatGPT Work, why Codex unexpectedly took off among non-developers inside OpenAI, and the company’s broader plan to bring useful agents from software engineers to knowledge workers and eventually everyone. We go deep on the shared agent harness behind Codex and ChatGPT Work, why OpenAI brought the experiences together without making them identical, and how persistent computers, artifacts, Sites, plugins, memory, and sub-agents are changing what people can delegate to AI. Akshay explains why some teams are replacing decks and spreadsheets with interactive websites, how agents can gather context across code, Slack, documents, and local files, and what OpenAI learned from personal-agent products like OpenClaw. Side note: also don’t miss Abhihek’s sandbox track keynote at AIE, which now powers a lot of the sandboxing for ChatGPT Work… and yes was also broken by an unreleased OpenAI model in the recent HuggingFace incident. Akshay also reflects on how AI is transforming product development itself: why more people will become generalists with a specialty, why ideas and taste become the bottlenecks when almost anyone can build, why LLMs still struggle to generate genuinely grounded new ideas, and why teams must distinguish increased motion from actual progress. We discuss: * Why Codex unexpectedly took off among non-developers inside OpenAI * Why employees felt like using Codex gave them a new superpower * The product insight that led OpenAI to build ChatGPT Work * Why Codex and ChatGPT Work share the same underlying agent harness * How their UX, Git visibility, artifacts, and sandboxing defaults differ * Why OpenAI merged its agent experiences instead of building separate products * How AI is blurring the boundaries between engineering, design, strategy, and operations * Why OpenAI wants the default model configuration to work for most users * When power users should use deeper reasoning, Ultra, or multi-agent modes * Artifacts, agentic spreadsheets, and creating high-fidelity work products * Why interactive Sites may replace decks and spreadsheets * The challenge of designing a simple interface for an agent that can build almost anything * Why users should retry tasks that models could not handle three or six months ago * How AI can gather context for performance reviews without replacing human judgment * The OpenAI automation that turns internal Slack and document activity into memes * What reaching ten million ChatGPT Work and Codex users means for the product * How OpenClaw inspired persistent environments, scheduled tasks, and personal agents * Using ChatGPT for financial planning, budgeting, workouts, meals, and household management * The design tradeoffs behind sub-agents and how much of their work users should see * ChatGPT memory, Chronicle, and long-term context * Why AI may make more people generalists with deep specialties * Why ideas and taste become more important when almost anyone can build * Why LLMs still struggle with the instruction “bring me new ideas” * Measuring productivity through quality at-bats instead of commits, tokens, or pull requests * The critical difference between AI-generated motion and meaningful progress Akshay Nathan * LinkedIn: https://www.linkedin.com/in/akshaynathan/ * X: https://x.com/akshaynathan_ Timestamps 00:00:00 Introduction and Bringing the Power of Code to Everyone 00:01:33 Joining OpenAI and Preserving a Startup Culture 00:02:40 What OpenAI Learned from Enterprise AI Adoption 00:05:28 Why OpenAI Built ChatGPT Work 00:07:17 Codex vs. ChatGPT Work and the Shared Agent Harness 00:12:07 Why OpenAI Merged Its Agent Experiences 00:16:24 Models, Reasoning Levels, and Choosing the Right Default 00:20:26 Artifacts, Agentic Spreadsheets, and Model–Product Collaboration 00:24:22 Why Sites Could Replace Decks and Spreadsheets 00:30:08 Designing an Agent That Can Build Almost Anything 00:34:28 From Developer Agents to Knowledge Work—and Everyone 00:36:07 Power-User Advice and AI-Assisted Performance Reviews 00:40:41 OpenAI’s Internal AI Memes and the Ten-Million-User Launch 00:44:39 OpenClaw, Personal Agents, and ChatGPT as an Operating System 00:50:24 Sub-Agents, Ultra Mode, and How Much Control Users Need 00:54:39 ChatGPT Memory, Personalization, and Chronicle 01:00:19 How AI Is Reshaping Product Development and Tech Roles 01:03:15 Ideas, Taste, and Why LLMs Struggle to Generate New Ideas 01:04:42 Measuring Productivity, Quality At-Bats, and Motion vs. Progress Transcript Introduction: Akshay Nathan, ChatGPT Work, and the No-Code Arc Swyx [00:00:00]: We’re here in the studio with Akshay from OpenAI. Welcome. Akshay Nathan [00:00:07]: Thank you. Swyx [00:00:08]: And with our trusty co-host, Vibhu. So you recently launched ChatGPT Work. You lead Core Product Engineering. It’s been a long journey, into all this. I find it very interesting that you started with no code or low code, with Walrus and Airtable. And to some extent, ChatGPT Work is like the super app of super apps of, well, here is the ultimate no code. You just write a prompt. Akshay Nathan [00:00:32]: Yeah. It’s funny how things come, full circle. I think for a long time in my career, I started my career working consumer fintech, but then after that, like, there’s this hypothesis that, the things that we were able to do with code, like, as engineers, like, if we could bring that to many more people in a more, accessible way, then that would be truly magical. We were working on a startup. It’s funny, like, before LLMs, before vision LLMs, on how to do automated testing with AI. It was just kinda jank, back then, but doing what we can, and then worked at Airtable for a while on the same thesis that, like, if we can bring a database or the primitives behind a database to people, that’d be really useful to them. But once LLMs came onto the scene, it became clear that, this was the missing piece, like, the missing technology required to, like, bring the magic of code to everyone without them having to know what’s going on underneath the hood. And so, like, I think this launch and a lot of the stuff that we’ve been up to is, like, the manifestation of that. From Walrus and Airtable to OpenAI Vibhu [00:01:33]: How was stuff when you joined? So you joined OpenAI 2023. Now we’ve got, so much more stuff, so ChatGPT, Codex app, ChatGPT Work. Have things changed? Joining OpenAI and What Hasn’t Changed Akshay Nathan [00:01:44]: I think the more interesting thing is how things haven’t changed. Like, one, I joined I remember when I joined, it was, like, five hundred people. One thing I was worried about was, like, I was looking for something, more early stage and, like, was it gonna feel startup enough? And I joined, and I was like, “This feels even more startup-y than I could ever imagine.” And, like, that really hasn’t changed even till now. I think the, like, level of, like, bottoms-up ambition and, like, the ability of anyone to, like, do anything or have an idea and ship it is really cool. But on the, like, mission side, I think what was really compelling to me is this mission of, bringing frontier intelligence to everyone. Like, building AGI and then bringing it to everyone. And, I think acknowledging back then that, like, that vision is gonna, not be a linear progression. Like, we’re probably gonna, like, try different products and have different things that succeed and don’t. But the vision has stayed the same, and the mission has stayed the same, and we’re starting to see the pieces, fall together, and that’s really cool. Enterprise Lessons: No One-Size-Fits-All AI Swyx [00:02:40]: You worked on Enterprise. What A lot of people never touch ChatGPT Enterprise. What is something that you learned from there that you’re bringing into your work now? Akshay Nathan [00:02:52]: I think how there’s no one-size-fits-all solution in Enterprise. I remember in the early days of ChatGPT Enterprise, like, when we talked to customers and, like, everyone. That was, like, when I think it was a year after ChatGPT was released, and everyone was so excited to bring, AI into their enterprise. And, there were all these teams being stood up. It was, like, the AI deployment team with, like, these enormous budgets. And if you asked anyone, like, what were they excited about? Like, what were they excited about solving? Like, at first, you’d get, like, kinda like the baseline answers of, like, “Yeah, we have all this context and data and all this stuff.” But then if you ask them, like, “What was, like, a discrete use case that, like, they want AI to enable in their workplace?” You get such a different, like, variance, like, explosion of, different types of answers. And it’s interesting, like, you using, like, these models and these products, you have this box, and you can say anything to it, which is the magic. But it’on the flip side, it also means that, like, you don’t know what to do with it. And in Enterprise, I think a big part of that is, like, meeting the users where they are, like, what use case were they trying to solve, and then teaching them how they can use AI to, like, gain leverage there. Swyx [00:03:56]: Do you meaningfully differentiate that from forward-deployed engineering? Akshay Nathan [00:04:01]: I think there is the go-to-market side of it and then there is the product side of it. I think you need someone on the product side. And I think, like, however good we get at FDE motion, like, I think at the end of the day, if we have a user who’s, like, looking at their computer or looking at their phone, like, it’s our job in the product to, like, be enabling them and showing them where to go. So we’re really excited about that. Vibhu [00:04:24]: Do you think there’s been changes, over the past three years of adoption? So there have been, step function changes. You have reasoning models and whatnot. Is there still the same problems of Enterprise has black box, don’t know what to do with it, or have things changed? Adoption, Agents, and the Next 10x Market Akshay Nathan [00:04:39]: We’re seeing now that, like, there’s this huge uptake, right? Everyone is extremely excited about it. It feels like, many people are, millions, hundreds of millions of people are using ChatGPT. They understand, like, how generally to work with AI. But then, like, every time, like, a new capability gets unlocked, so now, like, we’re seeing with agents, like, there is probably a contingent of, like, early adopters still who, truly get it, who are like, “ we you can do anything. You just have to make sure the right context is there, it’s connected to the right tools, and that you are supervising it, but, like, anything is possible.” But then there’s, like, this, like, 10x or 100x bigger market where, like, they don’t yet get that, or they don’t yet see that. And so I think that’s the next stage here. So to answer your question, like, I think the adoption is there and growing fast, but I think the opportunity is, like, far bigger than that. That’s where we wanna play, especially with ChatGPT Work. ChatGPT Work, Codex, and the Super App Merge Swyx [00:05:27]: Yeah. well, let’s, let’s skip ahead to ChatGPT Work. only, like, a month ago or so, announced. what was the decision process that led into it? there was this, overall merging of the super app. Is that what we’re officially calling it? you deprecated the browser as well. Just, summarize your last, like, couple months of working on this thing. Akshay Nathan [00:05:50]: Yeah. It feels like forever now, but it’s only been a few months. I think maybe the one, impetus that, like- Is most salient is when we release Codex, or even internally had Codex, like, it was really surprising to us, I think we recently put out some stats on this, that there was this, like, real inflection of, like, adoption among non-developers at OpenAI. And, I, through this product development process, like, would go to, like, these UXR sessions to talk to people internally. And the thing that stuck out to me is, like, one, like, you go talk to, like, strategic finance or marketing or whatever, and they’re all using Codex for, their use cases. That part’s cool, but the thing that really stuck out to me is how proud people were that they were using Codex. Like, how, like Swyx [00:06:34]: It’s like, “I’m not supposed to be using it, but I am.” Akshay Nathan [00:06:36]: It was that. It was, like, that they were, early to this, like, new thing, but it was also this thing of, like, they felt like they had a superpower, right? And, what we recognized then is that, like, the power of Codex, the power of agents, like, we already had this massive distribution base of people who have, come to know and love ChatGPT. Like, how do we show that to them? Like, how do we bring it to them? Which is, like, a hard product problem, and it’s, like, a tricky thing, right? There’s many ways you can go about it. And so that’s what we called the Merge and the Super App over time, and ultimately launched it in ChatGPT Work, is how do we do that? But it came from that initial realization that, like, the power was not only for developers, like, much earlier than probably even we thought. Like, it could be extended to everyone. Swyx [00:07:17]: How do you see the products differently? So, like, who is it for, right? So Codex started out even CLI, then app. Now there’s a merge of ChatGPT Codex and ChatGPT Work, so is it the opening for the average user, for enterprise, for work? How do you position it? Akshay Nathan [00:07:36]: I think we want to get it to position it for if you’re doing work-related things, for lack of a better word, right? Who ChatGPT Work Is For Akshay Nathan [00:07:42]: I think productivity is what, like, the pillar that I support. Like, that’s the name of the team. And the reason for that, the reason we call it productivity and not, like, enterprise or, like, work or something like that, is because there’s also personal productivity, right? And, like, I think ChatGPT Work is I’ve seen people do things in their personal lives that you wouldn’t classify as, like, work technically, but, like, these agents are, super capable for. Like, one recent example that someone posted about, on our Slack is, like, someone had, like, a missed package, like they didn’t receive it, and then they got, like, the picture of it, from Amazon or whoever the courier was, and they, like, asked ChatGPT Work to, like, find out where that package is. And, like, the agent, is extremely tenacious and, like, took the image and, like, looked at a bunch of, like, listings around their neighborhood and figured out exactly the apartment complex in which the package was, like, gave them some information. And so, like, I think there’s all these things that, like, you, work-related or productivity-related things, I think that’s what we want the product to be. You asked about Codex. I think we think Codex is, a durable brand, but we have a principle that, like, the user we don’t want a user to get stuck in a tab or an experience where they don’t get the power of the product. And so, like, everything that you can do, in the Codex portion of the product on desktop, you can do in ChatGPT Work and vice versa. But we made some opinionated product decisions on, like, how much of the Git state, if you’re in a Git repo, do we wanna expose to the end user? Or how much do we wanna make the experience of seeing the agents thinking, like, diff forward so that you get exposed to the diffs out of the box. And then, like, on the safety side, like, how do we wanna think about, like, sandboxing and making sure that we have the right defaults in one state versus the other? So, there’s, like, some opinions that go behind that, but we do want We don’t want the user to need to choose which experience they’re in. Swyx [00:09:26]: That is a good goal for AGI, right? Like, people don’t want, like, to hide to choose what version of AGI they want. They just want the AGI to decide for them. can I get an answer or, like It’s not super clear to me. Is the Codex harness and the ChatGPT Work harness the same? Is it just UI affordances, or are there prompt level or even deeper differences? Shared Harness, Different UX: Codex vs. Work Akshay Nathan [00:09:49]: So the harness is the same. The harness is shared. on In both of the products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plug-ins or computer use or artifacts. You get that power regardless of which experience you’re in. On the UX side, there’s opinionated takes that we have when you’re in Codex mode, what the UX should be how the UX should behave, and some stuff around the sandbox like I mentioned, but the underlying harness and capabilities should be the same. Swyx [00:10:16]: I’m just kinda curious. Maybe we can, -- Is there a query that we can run that would look different in the two modes? Akshay Nathan [00:10:23]: Yeah. I tried to create, like ask it to create, like, a retirement calculator spreadsheet or something, in both modes. And then in Codex mode, you might have to be in a repo for this, but you’ll see, like, the diffs of, like, the sheet that it’s creating and stuff like that, and the file edits. But in Work you won’t be able to see that. Swyx [00:10:42]: I think that’s, that’s super clear. And then also the other thing I wanted to dive into was your, the productivity team. what else is there? first of all, what are the top-level teams other than productivity? Isn’t productivity everything? Productivity Teams and Core Chat Akshay Nathan [00:10:55]: So Swyx [00:10:55]: Science? Akshay Nathan [00:10:55]: We have a team focused on ChatGPT. Like, the core chat experience, for consumer, which is like, not, I think all productivity. Like, there’People are using ChatGPT every day for search to, figure out how to write messages to loved ones, to think about, how to, like, learn a new topic, et cetera. And so there’s so much more inside to create images. And there’s so much more in chat that, the hundreds of millions of users are using that warrants, like, a very dedicated effort. And there’s teams focused on enterprise and infrastructure and API and stuff like that, so. Swyx [00:11:33]: I will bring it up. Retirement Calculator Demo and Git-First UX Swyx [00:11:34]: Yeah. So I have them both running. This is ChatGPT Work. There’s a Codex version here. I picked “Five Little Ducks” song, so this will take a while. Akshay Nathan [00:11:43]: Huh. Swyx [00:11:43]: I think we’ll just keep it in the background and, as they finish, we’ll look into some of the differences. Akshay Nathan [00:11:48]: Yeah. But immediately, I think if you flip back to the Codex version you’ll see that, Swyx [00:11:53]: That it assumes Akshay Nathan [00:11:54]: Like the Swyx [00:11:54]: It assumes Git. Yeah. Yeah. Akshay Nathan [00:11:56]: The, like, dynamic island assumes that you’re in a Git repo. And you might miss some stuff because some of it is, like, in the actual chain of thought with those changes and how we display that, but yeah. Swyx [00:12:07]: Is there an unintuitive like, is there a thing that you wanted to ship and then you got feedback, and you were like, “No, let’s not do it?” Like, what’s the thinking behind that? Why Merge the Experiences Akshay Nathan [00:12:14]: In, ChatGPT Work? Akshay Nathan [00:12:17]: I think one direction we could have gone with this is, like, keeping the experiences, like, completely separate. So it’s like, why Swyx [00:12:22]: Different apps. Akshay Nathan [00:12:23]: Exactly, like different apps or even in the same app, like different, completely different experiences. Like, why merge it all? Like, what is. Codex, people love. Like, why bring these products together? And I think the intuition here is that, like, all of our jobs are, like, changing dramatically with AI. Like, for, like, every few months, like, I feel like I wake up, and I’m, like, doing a completely different thing than I was doing a few months ago. And my hypothesis here is that, or I should say our hypothesis is that, like, part of what we’re, we’re building, this technology is giving people leverage. Like, the things, maybe it’s the more mundane parts of your job or parts that, like, if you were able to automate, you’d be able to share more ideas faster or whatever, like, you’re able to do now. And because of that, like, that might blur the lines between someone who’s, like, only writing code or creating strategy docs or, planning events or, helping with marketing or doing podcasts or whatever, right? And so, like, these things are gonna get blurred over time. And so, like, trying to draw a hard boundary based on, like, the who you are is gonna be, is gonna be tough. And, like, we should enable users to choose, but we shouldn’t box them in. And so a lot of the work that went in here, like, keeping the primitives the same, like for example, plugins are, like, unified across, this product and ChatGPT and the cloud, was because of that. It’s this thesis that, like, eventually things are gonna come together and we don’t wanna be Like, we wanna be prescriptive about when to be in either experience, but we don’t want to box anyone in. Swyx [00:13:45]: I wonder if there’s users who are very tuned to the old ChatGPT harness that is effectively now replaced by the Codex harness. I can’t imagine what that was, but maybe they’re more the more conversational side. Can you compare and contrast the two harnesses? ‘Cause only you’ve seen it. Akshay Nathan [00:14:02]: Yeah. I think ChatGPT, the existing harness, like, still exists today. Like, it exists in this app, Harness Engineering: ChatGPT vs. Codex Swyx [00:14:08]: The classic, right? Akshay Nathan [00:14:09]: The Vibhu [00:14:09]: You just start a new chat, and you don’t go under Work, right? Akshay Nathan [00:14:13]: Yeah. If you start Vibhu [00:14:13]: So Akshay Nathan [00:14:14]: A new chat and go to chat, then you’re, you’re talking to ChatGPT with the instant model. Vibhu [00:14:16]: Oh, we can technically do another. But on instant. Swyx [00:14:21]: Yeah. So this one’s not gonna code or it’s gonna be in line. It’s on a in line in a sandbox. Akshay Nathan [00:14:26]: It’ll Vibhu [00:14:27]: Oh, that’s cool Akshay Nathan [00:14:27]: We try to push you to go to Work if you’re creating a spreadsheet. Yeah, but this is Swyx [00:14:30]: And this is a router decision? Sorry. Is it a router decision? Akshay Nathan [00:14:34]: This is the decision that, the model is making, and then, like it sees that you’re able to. or you’re trying to do something that would be better served in Work mode. But I think your question was like, what are the advantages of, like, the chat, like ChatGPT chat harness? Swyx [00:14:48]: It’s more broadly, like, I wanna, do an oral history of harness engineering. Right? the ChatGPT harness lasted us from, let’s call it the ‘01 era, until now, and now it’s being replaced by the Codex harness effectively. And they’re, they’re overlapping somewhat, but I’m curious what changed if there is. Akshay Nathan [00:15:10]: My perspective on this is, like, there’s, there’s, there’s there’s like a constant process of, like, divergence, convergence, divergence, convergence. And in chat, like, many of the use cases I was talking about before, like, search or learning, I think we’re, we’re really optimizing for latency and optimizing for personality and, like, different things that, over time, like the product The reason people love ChatGPT is because we’ve been optimizing for those things and working on them for so long. Codex, what we learned was that, like, if you give the agent access to this infinitely flexible environment as a computer, it can do really powerful things. And so when we think about, like, okay, well, for knowledge work, like, what is which mode should we choose? It was like it felt more natural to us to bring that to this, like, computer environment and, maybe abstract some of the details of this computer away from users who might not be used to that, but, like, give them that same power. But ultimately, I think that we want the power in all places, right? We wanna meet people where they are. So I’m sure there’ll be work down the road in order to get things to be, equivalently capable in all scenarios. But it’s just a question of, like, what we’ve been focusing on the product on historically and what we’re focusing on now. Models, Defaults, and the Reasoning Slider Vibhu [00:16:24]: I think alongside that, outside of just harness and when to use Codex, ChatGPT, or Work, there’s also the new models you’ve released, right? any guidance there? So people love to min-max what to use, like only use Terra on high reasoning versus, for this, you wanna use Sol here, ignore all these Akshay Nathan [00:16:44]: There’s 32 options. Vibhu [00:16:46]: But, that being said, for people that are expanding, so, productivity trying stuff for work that don’t have the breakdown of what all this is what’s, what’s the advice, right? Akshay Nathan [00:16:59]: Well, I think before the advice, like the first thing is, like, none of this would be possible without these models. Like, the, I think you asked earlier, like, what was, like, the inspiration for work and, like, early on, like I mentioned, like, what we were seeing with Codex, but that was also because the models were getting infinitely more capable. That’s happening again. I think it’s like another step function jump now. And to answer the question on advice, like we want this default to be the best possible. Like, we wanna be opinionated about the default, and so we’ve we’ve chosen a default that we think is gonna be the best for everyone. And, we have for power users options under the hood. We could One could argue that there might be too many right now, and we’re, working on simplifying it. But you can extend, the reasoning level, and you can change between the different model classes if you need to, but the default should be the best for most use cases. So my advice to most people would be to stick to that. And then, if you reach a situation in which you think that you could, you wanna try, a different configuration, if you’re not seeing either the efficiency on the cost side or the quality on the intelligence side, then you can change the defaults and see if you can get something better. But we think that the default should be good enough. Swyx [00:18:09]: I have, I’m just gonna run something by you since you have way more experience than me. I’ve recently been doing Sol Lite but with goal, with the idea that the goal augments the reasoning effort, but with more terminations and turns. Swyx [00:18:24]: Is that a good way to think about it as opposed to Sol Ultra or Sol, Extra High? Akshay Nathan [00:18:29]: Yeah. It’s hard to say because Swyx [00:18:31]: Yeah. It’s like an interaction effect. Akshay Nathan [00:18:33]: exactly. It’s like there’s a preference on, for you as an individual, like how do you like to collaborate with the models? Like how many of those like terminations, as you call them, do you want where, you can steer or make sure that it’s doing the right thing? Akshay Nathan [00:18:46]: I think generally people should try whatever works for them. I think that like using Ultra or the like multi-agent setups are best for like when you have like tasks that are either incredibly complicated, like open explorations or very paralyzable. I think even for tasks using goal, I think is best for tasks that you’ll be able to make consistent progress in a way that’s verifiable over time. But I think for most tasks, they don’t fall into either of those buckets. And so like at least when they’re starting, and so that’s why I think the best first step is like trying it with the default configuration and then seeing like where you wanna go from there. Swyx [00:19:29]: Right. You guys worked on a slider, which is super helpful for reducing the amount of panic. Vibhu [00:19:36]: It’s nice on mobile at least. There’s a nice slider there. Swyx [00:19:38]: It’s nicer. Vibhu [00:19:39]: I haven’t tried it. Swyx [00:19:40]: So you have the advanced view there, but if you click advanced view. Yeah. Vibhu [00:19:44]: Ooh, it’s just a nice slider. Yeah. Swyx [00:19:46]: Very pretty, very colorful. Akshay Nathan [00:19:48]: Yeah. The idea was here was like reduce it to like one dimension even though there’s multiple dimensions, right? Try to project it onto a single dimension for the user. Like, something from that represents like, speed and efficiency on one side and then like quality and thoroughness on the other side. Artifacts, Spreadsheets, and the Work Launch Swyx [00:20:04]: I am just puzzled that it uses Sol so much, like the lower Vibhu [00:20:07]: No Swyx [00:20:07]: Grounds I would’ve used Vibhu [00:20:08]: I think the slider, if I’m not mistaken, is Swyx [00:20:09]: Terra. Vibhu [00:20:10]: Oh, it is. Swyx [00:20:11]: Yeah. See? So they preset Terra to only be the light one. But like I think a lot of people would more people should use Terra. One, because Sol keeps running out of capacity. Vibhu [00:20:22]: I’m the reason. Here’s ten minutes of our Swyx [00:20:24]: There you go Vibhu [00:20:25]: Retirement calculator. Swyx [00:20:26]: Oh, that’s the Excel thing working for you. Vibhu [00:20:28]: This is, Swyx [00:20:28]: Oh my God. Look at that Vibhu [00:20:28]: This is work, and then Codex is still cooking, so we’ll get back into it. I think it’ll be interesting to see the thought process, the reasoning, and also, this is eight minutes on work. Codex is still cooking. Swyx [00:20:41]: Yeah. And by the way, so I’ve, do Gabriel Chua? He’s part of the OpenAI Singapore team. He showed me this, and I was like pretty shocked that this looks like Excel. It edits Excel files. You never paid an Excel license, right? Like, but somehow this is like workable and it’s agentic Excel. Akshay Nathan [00:21:01]: Yeah. one of the big like pushes that we made for this launch was like artifacts, right? Akshay Nathan [00:21:05]: Like both on the model side, like I think if you compare this with GPT-5.5 and GPT-5.4 before that, you’ll see that there’s been pretty dramatic improvements in the quality of these artifacts and then also on the product side. Vibhu [00:21:16]: The UX side is also crazy, like hosted sites and whatnot. No longer needing to host your own little webpage, like it Swyx [00:21:23]: Oh, I have a story about that. I can do, a separate thing. I’ll need to take the visuals here, but we-we’ll, we’ll cut to that later. Was there co-training, because you were moving making this big move and you launched GPT-5.6 on the same day as ChatGPT Work? Was there influence between the model training teams and the harness teams, or did they did the launch dates just happen to line up the same day? Akshay Nathan [00:21:46]: I think the we collaborate heavily with the research teams, and I think that’s like one of the most magical parts of the job, like the most fun parts of the job. But yeah, just using artifacts as an example. Like, a lot of what you’re seeing, like underneath the hood, there’s a lot of work that went into making sure that like, we had the right infra to be able to train the models to get better at this. And then on the product side, like had the right experience for users to be able to collaborate with the model on an artifact like this. In fact, like this whole viewer, like the intuition here is that like, it’s not necessarily that you wouldn’t need an Excel license. This is stage one, right? Like, this is probably not what you meant when you’re like making a retirement calculator. Vibhu [00:22:24]: Yeah, you can iterate very easily. Yeah. Akshay Nathan [00:22:24]: You wanna iterate and like when you’re seeing it, and if this thing is high fidelity to like what you would see in or what your coworkers would see if you were to send this to Sean, like that I think makes it so easier and makes you trust the product in terms of iteration. Vibhu [00:22:39]: When you say coworkers would see, do you see a multiplayer, multi-team collaboration with artifacts? Any things you guys think about that? Multiplayer Artifacts and Collaboration Swyx [00:22:46]: You can already share it, right? Akshay Nathan [00:22:48]: Yeah. It’s inter It’s something that, we’re actively thinking about. one thing that, we’ve noticed internally without talking too much about the roadmap is that like there’s many times when someone will ping me about something, and I will ask ChatGPT Work the question, and then I’ll ping them back the answer. Akshay Nathan [00:23:04]: And then I’ll be thinking like Vibhu [00:23:04]: Like the simplest would be, the three of us are just all on one hosted. Akshay Nathan [00:23:07]: Exactly. And I’ll think about like was I required in this loop or and then maybe it was, rephrase like what they were asking or pulled from certain context or whatever. But like, when I gave them back the answer, that process was also lossy, right? Like I gave them just like my interpretation of what ChatGPT Work cooked up. But like underneath the hood, there’s so much context like in the rollout and stuff that could be interesting. Vibhu [00:23:28]: Yeah, it’s Swyx [00:23:28]: So like the answer was preemptively respond to every inbound request? Akshay Nathan [00:23:33]: No, it was just like literally like this is what I do sometimes as my job. Swyx [00:23:36]: I know you copy-paste and then you’re just a message forwarding service Akshay Nathan [00:23:39]: Yeah. Yeah, exactly Swyx [00:23:39]: From AI to AI. Vibhu [00:23:40]: But I think it’s interesting, right? It helps people understand the capability of what you can ask and delegate that oftentimes people don’t realize until they try or someone shows you, and then you’re like, “Oh, okay. Okay, I see.” Swyx [00:23:52]: I think it’s als there’s also like a, light security issue, where like you’re the permissions layer. Like yes, I could query everything that you query, and I could get an automated response, but maybe I’m not supposed to see it. And that there’s no way I would know because I’m not supposed to know what I don’t know. Akshay Nathan [00:24:07]: Especially as like, with ChatGPT Work, we’re, we’re asking you to connect your plug-ins and, it’s pulling from your local files and stuff like that. Like the amount of context that the agent has access to is like- Deeply personal and like that’s something I think we need to preserve, so that’ll be definitely a challenge. Swyx [00:24:22]: There’s Excel, there’s PowerPoint, there’s Docs, the, grand trio of work. What other formats of work do you think about? like you worked on Airtable. Is there a future where there’s like OpenAI Airtable? Like what does that look like if you ever ended up doing it? Akshay Nathan [00:24:41]: It’s a really good question. I think, Formats of Work: Sites as Knowledge Artifacts Akshay Nathan [00:24:43]: one that you didn’t bring up was Sites, and I think that was Swyx [00:24:46]: Sites Akshay Nathan [00:24:46]: A core part of this launch. There’s one side of Sites that I think people commonly talk about, especially on Twitter and stuff or X, of like, this like prototyping tool. And like we saw that happen with this launch even. The model slider that you guys were referencing earlier, like that was developed almost fully in a Site. Like, the collaboration between design and engineering and product on that was like on a site where we play with, the affordance and figure out how it feels and all of that. But the other aspect that I think is a little bit less talked about is like Sites as like an artifact for knowledge work. I was talking to someone the other day who’s on like our corporate finance team, and like we were mentioning how like now when they have these reports that they’re, they’re working on as a team month to month, historically those things were in slide decks and in spreadsheets, and now they’re just in Sites. And like Sites is the mechanism that they collaborate across the team. And the reason is ‘cause it’s like, it’s like somewhat higher bandwidth. Like, at these tools like PowerPoint and Excel are like infinitely flexible, but at some point you reach the boundary of like either as a human you may not know how to use some feature or something, or the product itself doesn’t support it. But with a site you can do anything. You ask for anything and you can get that. once people see that magic, I think it’s been really valuable. Swyx [00:26:02]: Yeah, let me show you my case study. this involves all the hot topics including ChatGPT Work, but also GPT-5.6 token billionaires and token maxing and Sites and auto research. I’m a fan of this game called Strata. It’s, it’s like a little board game that you Sites, Auto Research, and Research Dashboards Swyx [00:26:17]: That you play with, physical blocks, that come on top of it like that. So over the weekend I took like thirty photos and just threw into ChatGPT. one point seven billion tokens later, out comes this site with a fully playable thing Akshay Nathan [00:26:32]: Wow Swyx [00:26:32]: With 3D, block placement and everything. Because it requires physical blocks and I needed friends to train on it so they can get better, so I can play against them. But also, I could also, do things like train an AI on it and that’s, that Akshay Nathan [00:26:45]: That’s your auto research Swyx [00:26:46]: That gets into auto research. So, you want to train your own AIs, and then make sure they self-play against, each other. I need to set both AIs. So this is AI versus AI, and they’re, they’re gonna self-play. the AIs start out bad and then you want to define a loss function and get good. I wasn’t gonna supervise all this. I was at, I was down in San Mateo, attending a conference. What I ended up doing was, auto researching and on this and creating benchmarks and that there was just way too many parameters for me to read. So I started asking it for a site, and it’s created this lab, panel. Where is there a, is there a shortcut for a site that is created? Akshay Nathan [00:27:28]: You should be able to go in the sidebar to Sites, top of the sidebar. The left sidebar. Swyx [00:27:33]: This one? Oh, left? Akshay Nathan [00:27:35]: Yeah. Just scroll all the way to the top. Swyx [00:27:36]: Oh. Oh, it says Sites. Oh, there you go. Yeah. Akshay Nathan [00:27:39]: Ooh. Swyx [00:27:40]: So it create, it creates the sites. I don’t, I don’t think this is, it is exactly what I wanted, but let me show you what it popped up, right? Like I think as a research artifact, it is very important to communicate, exactly, what is being done. Outputs this thing which I eventually started publishing. So I moved it off of Sites because I wanted more, database and infrastructure than Sites afforded me. But this is like a research output that you can start to mess with and like try to think about like what hyperparameters are you tuning for training AIs. And like I was trying to make like scaling laws and everything and doing all sorts of like game optimization stuff. And the fact that you can just throw this up as a research artifact, like I no longer need to read ChatGPT output. I read Site output. But then there’s also a huge sprawl. Like look at how long this thing is. There’s so many numbers. It is pretty overwhelming, so then I have to start pruning it from there. But, it’s an interesting transition from Markdown effectively that you’re putting out to, you’re putting out a whole functional site. Akshay Nathan [00:28:41]: I think Markdown just isn’t that optimal for people to read, right? Might as well just write HTML website and I don’t know. I think you can do a lot with customizing this, right? You have your skills that explain what you want. Like I noticed they’re quite verbose. I don’t need a lot of this information. Swyx [00:28:57]: It’s very verbose. Akshay Nathan [00:28:58]: So and then the nice thing of having a site side by side is, you just iterate on what you want and what you don’t, right? Swyx [00:29:05]: Yeah. I don’t know if, any that triggers any stories for you of how it’s run internally. Am I doing this right? Akshay Nathan [00:29:11]: Yeah. I think that this is like a workflow that we’re seeing like all different types of teams use, where like the canonical artifact that was previously a deck or something is now becoming a site. And like with a site you, because it’s just HTML, you can like. It’s infinitely flexible. And so, if you want to give more prominence to a certain thing that like in a slide deck would, feel like it was buried, like you can do that. You can have it be like the hero image, right? And so I think that like, people are starting to see that. There’s more work to be done to make these things like much more easier, easy to collaborate on. You mentioned that they’re very, they’re long and verbose, could be broken up. I’m sure that there’s still something to do there. Swyx [00:29:53]: They’re super long. Yeah. Akshay Nathan [00:29:54]: Yeah. But I think we’re starting to see that like there is this aspect of this is a really interesting, format, for people to use, that’s like much more flexible than what they ever had before. Swyx [00:30:07]: I think your job also comes becomes meta. You’re not designing the products. You’re designing a product to make products, and I’m curious how you manage that. Designing a Product That Makes Products Akshay Nathan [00:30:18]: I think one thing that we’ve been Like when we look at the UX, like that we’ve been thinking a lot about is how can we balance like simplicity with capability? Like if we’re designing a product, like you said, that like is made to make up build other things, right? You can build so many different things. But we can’t put that all in front of you because you’ll get overwhelmed. Vibhu [00:30:41]: Yes. Akshay Nathan [00:30:41]: And so we had similar problem or similar challenges even Chat-with ChatGPT, but especially now, like when there’s so much that can be done, I think the balance that we’re constantly trying to strike is like, how can we give the user enough of a UI surface where, they can be expressive, they can tell the agent what they need, they can verify that it’s using the right tools, it’s pulling from the right sources, et cetera, but then it gets out of the way. And then how can we build the right system such that we can show them instead of telling them what can be done? Because so much of this is gonna be like, how do they discover the next use case and the next one after that if they really want to be super powered by the AI. Games, Private Evals, and Show-Don’Tell Vibhu [00:31:19]: Yeah. It’s interesting. I feel like everyone also just has a different way to do it, right? I made a similar version of this same game. I didn’t take any pictures of board or rule game. I threw in at goal eighteen minutes, fifty-three seconds later, a lot of tokens later, I’ve got a similar version. not with all the auto research and whatnot, but Akshay Nathan [00:31:39]: You gotta do all the latest trends. Vibhu [00:31:40]: And yeah, I did it with, did it with Codex, not Work, but it’s interesting, right? Akshay Nathan [00:31:45]: Yeah. And this is GPT Image generating the pro avatars. Very good for game design. Like Vibhu [00:31:51]: And Akshay Nathan [00:31:52]: A lot of game designers were like really into GPT Image for assets. Vibhu [00:31:54]: I will say like the broader takeaway probably is the reason that we do this is more so just to test the tools, right? Like, this was also a test for GPT-5.6 came out. I had done the game on GPT-5.5, right? The ability for me to no longer need it to. I had to feed it the rules. It’s, it’s a pretty niche game. It couldn’t find how to do this on its own. Akshay Nathan [00:32:15]: Oh, yeah. Vibhu [00:32:15]: GPT-5.6 Akshay Nathan [00:32:16]: It is out-of-distribution, which is why I was also very keen on testing the GPT-5.6 capability. Vibhu [00:32:21]: But, this is just as work comes out, as new things come out, these are just our side ways to test things, right? Akshay Nathan [00:32:27]: Yeah. It’s some private eval. That is not this private. Vibhu [00:32:31]: But also valuable because now you can send this to your friends and I learned about this game through seeing this. Akshay Nathan [00:32:36]: It’s a hard game. He’s very good. Vibhu [00:32:39]: It’s good to when no one is competing with you. But yes, it’s a classic RL problem of like self-play, bootstrapping your game AI. yeah, you see how easily work becomes personal and personal becomes work because the thing I do for personal, it directly informs people I work with because I showed it to them. They were like, “Oh, you can do that with GPT?” Which like I imagine is the growth strategy. Akshay Nathan [00:33:02]: Yeah. The show not tell is a big piece that, I think we’ve we’re not still not fully cracked of like, showing people all the things that they can do with the product versus like trying to teach that to them through like, articles or onboarding or whatever. Akshay Nathan [00:33:18]: So meeting them in the moment. Vibhu [00:33:19]: It’s a career risk for me, because I used to be in developer relations, right? Where your job is to show, and then you’re like, “What do you mean? You don’t, you don’t need.” your job is to tell. And then. But the product people are like, “Well, we don’t need you if our product is intuitive enough.” So Akshay Nathan [00:33:37]: Yeah. that’s the magic of the models. So you can tailor the telling or the showing to like specifically what the user needs, like what they care about, what they’ve done in the past, exactly where they are on the adoption journey. So I think that’s like gonna be a super big opportunity. Vibhu [00:33:50]: Seems easier and easier now to tailor custom showing, right? People have different use cases. As much as you said you don’t wanna segment different people into different buckets, right? It’s also not that hard to for people that are in different categories. But the question, is you said your team is more broadly on. What was the term you used? Productivity? From Developers to Knowledge Work to Everyone Akshay Nathan [00:34:12]: Productivity. Vibhu [00:34:12]: Productivity. So how Akshay Nathan [00:34:12]: Which is now work. Vibhu [00:34:14]: Is it work? Is there another distribution that we’re not hitting? Is there a group of people that will have something different than ChatGPT, Codex or Work? Is there more that the mass isn’t targeting? Akshay Nathan [00:34:28]: I see it as like a sequencing, like. The vision is like bring useful agents to everyone. We started with like developers. Like developers historically are like early adopters that are willing to put up with more friction, set things up, et cetera. Like that’s where, Codex started. I think the next opportunity is like what we call general knowledge work, all the other functions around developers. I think when you go from developers to this segment, like there’s inherent challenges with like, this show not tell thing that we’re talking about, making the product more understandable, bringing in new capabilities that matter more for this cohort than matter for developers, things like artifacts, things like computer use, et cetera. And then I think like the same learnings, like similarly how we took the learnings from developers and brought it to, general knowledge work, the next stage will be like taking the learnings from general knowledge work and bringing it to everyone no matter what they’re doing in their lives. And we’re already seeing that a little bit. Like this game example that you have is, something that’s like on the border of like fun and personal life to, your professional life. I use ChatGPT Work full-time at home for everything, like for whatever I’m doing. I used it the other day to come up with a meal plan and like, save that on the like computer environment that it has and something that I can continue going back to. Like is everyone doing that yet? Probably not because the thing says work on it, but eventually, we wanna get people there. Vibhu [00:35:51]: ChatGPT life. Akshay Nathan [00:35:52]: Yeah, exactly. ChatGPT cooking. But I think there’s a lot of, there’s a lot of opportunity there, but I see it as like, we’re, we’re built we built a foundation in software engineering, and we’re gonna take the same learnings that we take from software engineering to knowledge work to everyone. Vibhu [00:36:07]: Do you have any power user advice? I feel like, there’s a group of people that will live it, use it for everything, stay on it twenty four-seven. And then there’s a bit of a gap between that crew and people that, okay, I use it for work. I use it occasionally. Sometimes I type questions. any advice, any learnings, anything you recommend or just, takeaways that you’ve found that help bridge that gap? Power User Advice: Push the Frontier of Imagination Akshay Nathan [00:36:30]: I think a couple things that I’ve seen is like, one, that it really helps to broaden your imagination of what’s possible, and this has been a learning even for me. Like, the technology has progressed so fast that, something that, like, even three months ago, like, no way the models can do this. Like, now it’s like, wow, it’s like it can. Like, Swyx [00:36:52]: Give an example Akshay Nathan [00:36:52]: We’re going through right now our, like, review cycle internally, and, people always talked about this as, like, a thing that the models are good at and like, there’s a cliché of like: Okay, like, no one wants to be writing reviews and, like, we just use AI to do it. But in all seriousness Swyx [00:37:09]: And it can evaluate it as well. Akshay Nathan [00:37:10]: Yeah, exactly. In all seriousness, before it was, like, just, like, slop and, like, I think it was helpful, but, not super productive. Now I’ve found that, like, the model can do a much better job than me, especially in this environment of, like, pulling context on, like, what people are up to, how they’ve like the things that they’ve done to make a difference, highlighting like, wins that they’ve had that, like, I might may not even have seen. It has access to, like, everything, right? Like the code, like, things that they’ve caught, reviews, Slack, everything. And so it’s, like, incredibly powerful in that domain and, like, just like six months ago, the last time we did this cycle, like, I didn’t even I tried using it, but it was not at all helpful. And this time it’s been, like, incredibly helpful and, like, so I think continuing to push the frontier of imagination of what’s possible, even if you tried something before, I think is maybe the my biggest piece of advice. The other, thing is, like, the more you put in, especially in this environment where, like, the model has access to everything on your computer or in ChatGPT Work, like you can create, artifacts over time and save them in your library and, like, the model will continue having access to those. Like, the more information you give it about whatever domain you’re in, whether it’s your life or your work, the more valuable it becomes, and it’ll become valuable in, like, ways that might surprise you. Like, it might pull from context in a way that, may be proactive and that you might not even have thought about. But it needs to have access to those, to that those tools or that context first. Reviews, Agentic Search, and Context Gathering Swyx [00:38:27]: One thing I just wanna talk about the review stuff because I’m still that’s a very sensitive thing and you’re, you’re a founder, you’ve managed people, you’ve hired people. As manager myself, I’m very reticent to put out any LLM-generated things especially when it comes to people, ‘cause it feels like you don’t care. Swyx [00:38:46]: Presumably at OpenAI, people are more open to being eval rated by GPT. But are there any unofficial rules around this? Like, what’s the etiquette? Akshay Nathan [00:38:57]: Oh, I think the etiquette is that, like, I would never write something via, like, well, solely via AI and, like, present it as, like, a review for someone. What I was talking about is more, like, gathering context. That’s the place where it’s incredibly helpful. Swyx [00:39:08]: So it’s just search. Akshay Nathan [00:39:09]: Yeah, exactly. Swyx [00:39:09]: It’s agentic search. Yeah. Akshay Nathan [00:39:10]: It’s like agentic search, but, that you can tailor and steer much more capably than you could before, ‘cause, like, the thing is it’s all there’s a flywheel happening, right? Because of Codex, people are able to do, and because of ChatGPT, people are able to do so much more now than ever before. And if you’re able to do so much more, it’s easy to miss things as well. And so, like, I think we need to use these same tools to keep up with all the impact that people are having and understand, where we can be helpful. Swyx [00:39:39]: I think the thing, like, I run a small company, so easy to search, but at the scale of OpenAI with the amount of messages that you guys put in Slack, do you think that it misses things? Remembering What Humans Miss Akshay Nathan [00:39:50]: Probably, but I think that I also miss things. Swyx [00:39:52]: Like, it doesn’t matter, right? Vibhu [00:39:53]: I think sometimes it’s Swyx [00:39:53]: Like it’s, as it needs to be human-level Akshay Nathan [00:39:54]: It’s all relative, right? Yeah. Vibhu [00:39:56]: Sometimes it’s nice when it finds things you wouldn’t, right? Like right now, my Codex system prompts, they’re set up in such a way that every project I have has a secret- separate, notes MD, and it just writes learnings to there. And then the global one can pull from all these. So sometimes it’ll be like: Oh, there’s this project you did like four months ago. Here’s a note that we had, and it randomly pulls it back into context that I would never do, I haven’t thought about. Vibhu [00:40:20]: And I’m like, okay, this is quite superhuman, right? Like, stuff that would. And, it’ll save like hours on chunking of stuff or find something that’s already been done. I’m like, as much as it might miss stuff, I would too, but it’s very useful when it finds stuff. And I have like a very, non-super engineered solution to this. It’s just marked down files that get pulled whenever they want. Akshay Nathan [00:40:41]: Yeah. I have a funny anecdote about this. Like, recently gearing up to this launch, the team has been, really cooking on it for a couple months, and over that time, like there’s so much conversation and chatter going on in Slack and Docs and elsewhere. And, one of the members of the team set up this, scheduled tasks, like automation to like look at everything that’s going on and, like, come up with the best memes and then post it in one of our shared channels. And like, there are two cool things about this. Like, the first is, like, I think the models are, over time, like starting to become like funny. Swyx [00:41:13]: Funny. Nice. Akshay Nathan [00:41:13]: Whereas like, a year ago, like that was not at all the case. The second is, it was what you were saying, like they find things that in surprising ways that you may not have thought of and like create connections that you may not have thought of. And that really helps with like the meme generation because then you can see something that, genuinely surprises you and, is funny in that way. So yeah, that’s like not like the most productive, use of this the technology, but it does it does uncover this, like this capability that’s emerging, which is just like to find information that you otherwise would not know of. Launch Momentum and the 10 Million User Milestone Swyx [00:41:43]: Talking about the launch, I think, I have pretty much said this is the most successful launch in a long time. I think even more successful personally than 5.0, and they’re announcing ten million users. Does it feel different? You’ve been through a lot of launches. Akshay Nathan [00:41:58]: I think it feels like a culmination. Well, I think two things. One, it feels like a culmination, like I was mentioning earlier, like this like vision mission that we’ve been on for a long time. Like I said, we saw the magic of Codex internally, and then we’re like extremely excited to bring this to many more people and to see it working, to like see us reach, the distribution goal, numbers that you mentioned, like I think that’s like huge and super exciting. The flip side of that is like, there’s so much more to do too. Like, that’s also really exciting. Like, ChatGPT as a whole, like the this product that, everyone almost equates to AI and like loves, has hundreds of millions of users. And so like ten million is really cool, but like we need to get this to everyone. Like, we need everyone to feel this magic. And so that’s the next step from here. But yeah, I think extremely pumped about how it’s going so far and the opportunities. Swyx [00:42:46]: Awesome. I did want to also Because I’ve, I’ve, I’ve been tracking the number closely, it transitioned at some point from just Codex users to Codex plus ChatGPT Work, because they’re same harness. The whole point is that you don’t, you can’t, count them separately. Do you have roughly a billion, ChatGPT users? Why did it just jump to one billion right away? Like, isn’t that the default on ChatGPT or no? Codex, ChatGPT Work, and the Developer Brand Akshay Nathan [00:43:11]: We don’t default you into ChatGPT Work if you’re on ChatGPT Swyx [00:43:14]: If you’re free. Yeah Akshay Nathan [00:43:15]: It’s also only available to paid users right now. And I think there’s like a process of, educating users of what is the value of this product, having them try it, learning from their feedback, and making it better over time. But the goal is to, get as many of the people who love ChatGPT today to like feel the power of ChatGPT Work. But I think it’ll be a journey. Swyx [00:43:36]: Yeah. And Codex will still be alive as a brand for the foreseeable future. And we’ll just toggle between them as needed for UI stuff. Akshay Nathan [00:43:44]: Yeah, I think it’s even stronger point than that. Like, I think we fully intend to like, treat developer. Like, developers have been, a core market for us for so long, and like there’s, there’s so much more that we can do to make Codex great specifically for, software development, and we’ll continue to do that. This doesn’t take away from that at all. If anything, it should increase the utility of something like Codex, because now you can move seamlessly between writing a diff to creating an artifact or, doing a search over your factor. Swyx [00:44:11]: I do wonder how much this terminology leaks to the non-technical user. Like, do they have to learn to say artifact if I want artifact? Or. Akshay Nathan [00:44:20]: It’s funny, like we call it artifacts internally ‘cause that’s what the teams call it. Swyx [00:44:23]: It’s nice. Yeah. Akshay Nathan [00:44:23]: But like externally, like no one says that, no one calls it an artifact. But I think that people like often, like describe things, whatever they’re used to, right? So if, ChatGPT Work is good at creating slides, they’ll say ChatGPT Work is good at creating slides, and that’s what we want. OpenClaw, Personal OS, and Persistent Computers Swyx [00:44:38]: One big Another, it’s July of twenty-six. One big thing that also happens in, for OpenAI was OpenClaw, and that’s I think a lot of people’s first time really maxing a agent for personal stuff, but also crossing over to work in essence same way. As far as I understand, OpenClaw is still independent, but did you go through your own OpenClaw moments? Were there any lessons you took from OpenClaw to Codex or back? Whatever. Akshay Nathan [00:45:06]: I think there’s a lot of inspiration. I did go through my own OpenClaw moment. I, Swyx [00:45:10]: Yeah, tell the story Akshay Nathan [00:45:10]: Me and my wife like set up an OpenClaw to like try to manage everything in our house. Not that there’s like a ton, but it was like quite useful. We gave it a calendar. It started, creating events for us and stuff. At some point, the laptop that we were running on, it died and never got a chance to pick it back up. But there was a lot of inspiration there, like, in ChatGPT Work, in web and mobile, like you get access to this like persistent computer environment where, you can store files, and those files stay around between sessions. And the idea is to be able to enable use cases like this. one of the members of our team uses ChatGPT Work for what they used OpenClaw from before, and then feel like it has like completely transitioned, which is like, workout planning and like meal tracking. which again, it’s like a work-related thing, right? It’s like not work necessarily, but it’s like in personal productivity space. But it has all the same primitives. So it has scheduled tasks. It has the ability to store files on a file system. It has the ability to like reference those things over time. And so you start to see the same types of use cases emerge, which has been really cool. Swyx [00:46:14]: Is there a point that ChatGPT Work completely replaces OpenClaw? they’re independent, so. Akshay Nathan [00:46:20]: Yeah, I’m, I’m not close to it, so I can’t speak to the OpenClaw roadmap, but I don’t think so. I think that there’s gonna be, there’s always a need for like this like incredible, like open source technology that team has built. And I think that we can draw inspiration, in the product and, ChatGPT, I think many more people have like heard about and used ChatGPT than have used OpenClaw. And if we can take the magic from OpenClaw and bring it to them, I think that’ll be a success. I think that like one thing on the ChatGPT Work side that we feel strongly about is that like the core experience is that you come to this product and you have a conversation, start a session, whatever you wanna call it, with this agent. And the magic of the product is that you can do anything in that moment. And we would like to create a product where you don’t have to click a button or to go to a different place, whatever, and you can get whatever functionality exists in, your finances app or where or any other product like in this one place. And so that’s the goal. It’s like it we want an extensible system with plugins where you can connect to the tools that you need in order to be able to accomplish like a financial task, where you can, if you’re doing like science work, like we have an ability to like extend the system in such that you can like write the tech and it performs well. There’ll always be like products that we support that are best in class at those things, but we want as much of the magic as possible in that core experience. Swyx [00:47:45]: Yeah. Do you think that you can do everything you used to do with Wealthfront in ChatGPT Finance? Finance, Data Access, and Centralized Context Akshay Nathan [00:47:50]: I tried it. like ChatGPT doesn’t yet custody, cash and assets for me. So that part, no, not yet. But I, there was like a whole component of like retirement planning and, like financial planning and budgeting and stuff that, we were looking into when I was there. And like with the finances plugin, like that’s all possible with ChatGPT today. So, I feel like at least that component’s replaced for me. Swyx [00:48:17]: I haven’t really plugged it in yet. I’m somewhat scared to look at the answer. Like that’s honestly like the same reason for health and finances. Like I’m like, no. Akshay Nathan [00:48:27]: It’s really good. It’s really cool how we were talking about like the agentic search aspect a little bit earlier, but like, it’s really cool how like, in conventional UX, like if the more power you wanna give to a user, the more like knobs and bells and whistles you need to add. Like, for like these finance and budgeting apps, like there’s always like a bunch of the different filters and like search bars and stuff like that. But like now, like with the right Vibhu [00:48:48]: Connect-connectivity to the right data, you can have whatever you want. You can ask any question you want and into that box and get the answer, and I think that’s super powerful. Akshay Nathan [00:48:57]: I think it’s also nice to just have it centralized in one space, right? You have different health apps. I have one for a smart scale, a watch, all these different things. It’s just nice to centrally co-locate it. Vibhu [00:49:08]: Which is, part of the whole thing of OpenClaw, right? Like that you would have, personal OS, which presumably ChatGPT wants to become. I do think that just relying on, like, just-in-time pulling of data for, let’s say, through via MCP, CLI, API, whatever you do, still not enough. Like I come from a bit of a data engineering background, like you still want like a data warehouse or some caching or semantic layer. do you feel that or do you already have that? Akshay Nathan [00:49:40]: I can’t speak to like all the details on how everything works, but I think it depends on the access pattern, right? Like if you want an answer immediately, then yes, it’s very difficult to do that if you need to pull from all of these sources. But a lot of the like use cases that we wanna enable in ChatGPT Work aren’t necessarily something that you need immediately. It’s more like a task that you want the agent to go and do, and that’s gonna take a certain amount of time. And, with things like programmatic tool calling and stuff now, like some of that time and sub-agents and stuff, like some of that is also parallelizable. And so it’s possible I think it’s very possible that there’s a, the ceiling on what can be done, with MCPs and like calling out to these third-party services has been raised substantially. So we’re really excited about that. Sub-Agents, Ultra, and Product Design Tradeoffs Vibhu [00:50:23]: You mentioned sub-agents. I gotta double-click on that. Ultra is a new mode. You have special affordances in ChatGPT itself to show off the agents. Can’t really do much with them, to be honest. Like just watch. what have been, what have been your experiences, any design issues that you would call out to other builders building with sub-agents? Akshay Nathan [00:50:45]: I think it’s goes back to the balance that I was raising earlier about like, showing builders the power of the tool, but also creating enough of an abstraction to not overwhelm them. I think with sub-agents, the thing that we wanted to show is that you can take a task that, has many parallel tracks or, is complicated in a way that, sub-agents can handle, and this product is for you. Like, the model can accomplish those goals or try to accomplish those goals. And so like that’s the point of like showing them in the product and that’s where we-we’ve gone with the design. There’s another, iteration of this where like you can see exactly what they’re doing and things like that, which I think is like, could converge on like overwhelming, with information. And so this is like the deliberate trade-off that we made for now. Vibhu [00:51:33]: You do display quite a lot of transcripts. Akshay Nathan [00:51:35]: Right. Right. Vibhu [00:51:36]: Or do you Akshay Nathan [00:51:36]: I think it’s hidden by default though, right? Vibhu [00:51:37]: Do you want to display more than that? Akshay Nathan [00:51:38]: No, it’s hidden by default. Yeah. Vibhu [00:51:39]: Some people could want more. So I’m one of those people that will throw a lot of stuff at goal, and pretty much every goal I’ll tell it to use sub-agents. Seems redundant, right? But every time I’m like, “Okay, use sub-agents where possible.” And I have a lot of people, a lot of friends that recommend and do the same. Whereas I’ll sometimes talk to people that are like, “Okay, this is where I want you to use sub-agents for this sub-task,” and I’m sure they would appreciate seeing into how they’re being used. For me, it’s primarily like two things, right? One is net time efficiency, so span out across sub-agents. Two is probably cost, right? Vibhu [00:52:15]: Don’t use big, expensive model. Offload to a lot of smaller, cheaper models. And some people want that level of control. So if you have repetition in what you’re doing, right? Say I want something built where I want it to consistently do this every day, I might wanna go in and fine-tune sub-agents here, sub-agents there. So you can see both, but I think if I’m not mistaken, it’s hidden by default. There’s a dropdown that goes a lot where I’m like, okay I’m just gonna keep, using. Akshay Nathan [00:52:41]: Oh, you can change the model that they use. Vibhu [00:52:42]: I know I tell them to be steered. I’ll say my I know Anthropic offers this in Cloud Code. You can tell Fable to use Sonnet or Opus to use Sonnet as sub-agent, so pretty trivial thing. You tell it to span out sub-agents with Sonnet, it’s cheaper, faster. I would assume if it’s not there, it could be built there. But I think there’s a side of Akshay Nathan [00:53:02]: It’s too many toggles. Vibhu [00:53:04]: It’s not a toggle. It’s just, you tell it in chat. Akshay Nathan [00:53:07]: You’re prompting it. Yeah. Vibhu [00:53:07]: The way I do it is prompt it, right? And I think this is something that gets abstracted unless it’s something you built for repetition, right? So if I’m building something, say that’s, podcast prep, right? Research into people, do a very deep extensive research, that I might wanna configure to cheaper, faster model just for web search, right? I can see a world in which you want both. I think the default is pretty good right now, where it’s hidden, but you can drop down and get some more info into what’s done. Vibhu [00:53:34]: I know people talked a lot about it on GPT-5.6’s launch. this thing loves to use a lot of sub-agents and causes the ChatGPT app to just crash because it’s so processor-heavy. But, Akshay Nathan [00:53:47]: For what it’s worth, that’s not my experience. Yeah, I haven’t had a crash from sub-agents. Vibhu [00:53:52]: I haven’t either. I have We both have big laptops. But I know people brought it up. There was a topic of discussion that we didn’t see the same, but it is another vibe eval, right? People are like, “Okay, the amount of sub-agents Sol is wanting is crazy.” And I’m like, “I think this is okay. I think it’s good.” But just stuff people bring up. Akshay Nathan [00:54:12]: I think when we launched the product too, we weren’t as opinion about like who is Ultra for and like when should they be using it. And since then we’ve made some changes to like, require you to turn it on and find it in the advanced setting ‘cause that’s who it is for. It’s for like power users who understand what’s gonna happen because it also, depending on your use case, can use more of your limits as well. Vibhu [00:54:33]: Yes. Akshay Nathan [00:54:33]: So that’s where I think a lot of the feedback was coming from. Vibhu [00:54:36]: It’s okay. Reset the limits. Always reset the limits. Akshay Nathan [00:54:39]: Well, it’s, today we’re resetting because of this. I wanna change topics to one last piece of the harness, memory. A lot of people are commenting on memory recently. ChatGPT’s new memory system used to suck, it’s not very good. And then this guy also the same thing, and Samir, who you presumably work with Memory, Chronicle, and Personalized Context Akshay Nathan [00:54:55]: Talking about memory. What can you say there? I think that, Samir and the team have made a ton of and then the research teams have made a ton of, updates and improvements over time. I think when I talk to friends, family members about what they love about ChatGPT, like the fact that it knows them, that they feel like their ChatGPT is their ChatGPT, I think comes up probably number one. In ChatGPT Work, in the Cloud, like by default, all conversations like inherit from your ChatGPT memory, so you’ll know they’ll know context about you, and they’ll also be able to write back to this memory. Vibhu [00:55:27]: With it, like a small text write. Like you tell me when you’re writing, right? Is it Akshay Nathan [00:55:31]: No, it’s part of the same like memory V3 system that we launched. Vibhu [00:55:36]: Yeah, Memory V3, yeah. Akshay Nathan [00:55:37]: So I think that’s been really powerful because, going from ChatGPT to ChatGPT Work feels like an extension of what I’ve already been doing with the product for sometimes many years. So that’s been awesome, and it’s awesome to see that like people are recognizing the improvements here. Vibhu [00:55:51]: Is there So it’s a retrieval problem, right? Like, are you retrieving the right things? Are you over-focusing on the wrong things? Is there like a more false positive or false negative, if that makes sense? Like, what’s the bigger problem? Akshay Nathan [00:56:05]: So I don’t work on memory directly so it’s hard to say what the bigger problem is with like certainty. But I think you’re right. I think that like, the there’s two sides of it. It’s like, making sure it knows things about you, but then also having the EQ to like bring those things up at the right moments proactively or surprising you in ways that are positive, not negative. Akshay Nathan [00:56:21]: So I think it’s a very challenging problem, but something that I think we feel very there’s a huge opportunity to get right, which is like why we’ve made like big investments in it. Vibhu [00:56:29]: How do you see the side of, okay, when you’re building ChatGPT for work different than the regular chat app, different than Codex, managing memory across different projects, collaboration and whatnot, how do you see the side of what’s separate from the harness, right? So if I have four threads on one project any learnings on how to build memory systems there? For background as well, to steer it a bit, is when you do chat style applications, I’d say you have a lot of one-offs, right? Vibhu [00:56:58]: When you switch to work it might be something you’re doing for a month, something you do a lot, right? Now, as I add more sessions, there’s a lot more than just single-threaded, right? Vibhu [00:57:08]: And there might be memory there. Akshay Nathan [00:57:10]: I think first I challenge that like the depth of the memory or the like value of it is like fundamentally different across chat and work. Like it is true that like, there are a lot of like shorter sessions on chat, but I think, the ChatGPT, the product has had like a ton of longevity, in, as long as this technology has been around and people use it for work-related, like productivity-related things already today. And so I think we found that there’s a lot of value. I found this my personal usage, like all these one-offs add up over time into something like quite durable and like quite a good representation of who I am. I know like from time to time, something will go viral on X about like, ChatGPT telling you everything it knows about you, and people are always surprised like how deep that is. Vibhu [00:57:55]: The fun roast me? Akshay Nathan [00:57:57]: Exactly. So like, I think like the That’s all to say that like I think there’s a lot of depth there in the existing, ChatGPT product, and so that’s why I think we think it’s valuable to bring into the work product. But the other reason I brought that up is because I think like hopefully we can use some of the same fundamental primitives and systems to extend memory here as well, and I know this is something that the team that focuses on this is like working through right now. Vibhu [00:58:20]: I wanted to bring up one element of memory, which I honestly don’t really use much, and I’m curious if you do: Chronicle, which was, is up on screen right now. It’s a super memory or like what is it? Akshay Nathan [00:58:33]: I think the idea is that like it can learn from, how you’re using your computer and like it’s another input source, into memory. And, I think it’s, experimental right now and something that like isn’t default off. But I’d recommend that you try it. I think that it’s like quite interesting how It goes back to a conversation we were having earlier on like, you were asking like, “Does it Can ChatGPT miss things?” Like does it, on Slack, when it’s searching, does it miss things? ‘Cause there’s such a volume of stuff, right? And like it’I, you can ask the same question about like everything that you’re doing on your computer. Like, is it gonna know everything that you’re doing? Is it gonna capture the intent and stuff like that? Probably not, but like it probably will find things that you might not know about. And then if it can surface those to you in relevant times, in proactive ways, like when you’re doing tasks, and I found at least that it can be quite helpful. So it’s worth trying. Vibhu [00:59:24]: So mostly for insights and longer term. Akshay Nathan [00:59:27]: Yeah, exactly. Like insights and it builds context that makes, that can make you more productive on certain tasks. But it’s, it’s hard to describe without feeling it. Vibhu [00:59:37]: I will say you can feel it pretty well. Like the idea of what they’re saying here, right? Just check through my memories or check through my logs and add skills. Pretty underrated, right? Akshay Nathan [00:59:48]: But that’s automations. You can repeat that using a cron job. Checking through your memories and creating skills. But I think the creation of the memories from Chronicle itself is like what’s different. It’s like you have much deeper memories because you have Chronicle on. Vibhu [01:00:01]: It’s there. I don’t use it much, but maybe I just, I need more examples. I imagine you guys use a lot of it internally, so I’m always fishing for use cases. Akshay Nathan [01:00:10]: I would just try turning it on and then like Vibhu [01:00:13]: It just auto works? Like it Akshay Nathan [01:00:14]: Yeah, and seeing like where it might start helping you. I think you’d be surprised. Vibhu [01:00:18]: Yeah. Amazing. I think that was, about it in terms of like the overall, coverage of ChatGPT Work. I think there’s been a lot of like good progress and discussion on building and all these things. There’s a lot of like ex-founders in the community, in OpenAI as well. Do you think that things have changed a lot? like your overall reflection of building, pre-AI and post-AI. Akshay Nathan [01:00:44]: I think things have changed a ton. I think it’s like super exciting to see how quickly you can go to, from idea to something real today. whereas like even before, like I think, five, 10 years ago, like it’s fast if you were scrappy and, like, willing to build the minimal viable thing. But, like, now the extent of what you can build is, like, much broader. And I think that also, like, what we’ve seen internally building is, like, that gives you an opportunity to validate much more quickly, to talk to users, to talk to internal doctors, et cetera, and, like, make sure you’re on the right track. And, like, that loop I think has been has become more closed than ever before, and that’s, like, a win for product development. I think it’s a win for consumers and users too because ideally that means they’re getting much more better much better products out the gate. Building Before and After AI Vibhu [01:01:32]: Does it mean your teams are smaller? Akshay Nathan [01:01:33]: I think there’s much more to do now. So I think people can accomplish more individually or in a small team than they were that would require more people than before. But there’s, at the same time, there’s also more to do, so I think the teams are much more ambitious. Vibhu [01:01:50]: Have you seen any changes in scopes of roles and building teams and how we used to have teams, say, a few years ago versus what ideal teams look like now? Akshay Nathan [01:01:58]: I think we’ve seen a blurring in the lines between, like, the typical product development functions, like between, like, EM/PM, engineer, designer, et cetera. Like Vibhu [01:02:08]: Yeah, I wanna bring up this quote. There will be, only four jobs left in tech. There’s AI slop cannon, the people who just, like, they’ll burn a bunch of tokens. And then there is SRE, the people who. people who are more responsible. There’s grown-ups who sell things, and then there’s hot people. Akshay Nathan [01:02:27]: This is an interesting take. I think my suspicion is that there’s everything everyone will be, like, shaped in a way, in that, like, AI will enable everyone to become a generalist. Like, things that, like, I never would be able to, like, come up with a design before and, like, even now, like, I don’t have maybe, like, the visual taste required, but I can iterate on something with the help of AI. But then people will have a specialty, and that’s, like, the straight line in the T or the upward line in the T. And so, like, you can have a specialty that you’re interested in. With the help of AI, you can go deeper and become better at over time, but then you’ll also be a generalist. And so with that foundation, the way you can accomplish is, like, almost limitless. Team Shape, Shaped Builders, and Taste Vibhu [01:03:07]: What are you bottlenecked by in terms of specialties? Like, do you need more designers? Do you need more slop cannons? Do you need more hot people? Akshay Nathan [01:03:15]: I think the bottleneck some becomes, like, ideas and taste. I think because anyone can build now, I think, it really is the era of, like, bottoms-up ambition. And because there’s so much to be built, like, you’re always gonna be bottlenecked by, the amount of ideas and amount of things that you’re doing at any given time. Vibhu [01:03:37]: Do you think models help solve that? Akshay Nathan [01:03:39]: Models? Vibhu [01:03:40]: Yeah. I have the example of, like, I have a front-end design skill that’s like, they give me four drastically different examples of what this looks like. Sure, it burns a lot of tokens, but. And then I’ll mostly just condense down, “Okay, I like this part. I like this part. Let’s draw these together.” And it’s like, yeah, I had a vision, but, like, I don’t know. Akshay Nathan [01:04:01]: I would say that the one automation that I would love to work and it doesn’t work is bring me new ideas, right? somehow LLMs are just not it. One interesting part about ideas is, like, they’re not, like, in a vacuum. It’s, like, not. They usually come from somewhere and, like, in product development, like, they’re coming from talking to users or reacting to, friction that you’re seeing or feedback, building on some foundation that you already had planned out before, whatever. And so I think that’s where, like, I think there will always be value in these, like, generalists that we talked about, like, closing that loop and then having coming up with those ideas that are grounded in that feedback or talking to users, whatever it is. Defining and Measuring Productivity Vibhu [01:04:41]: Cool. You were gonna. You lead the productivity team. How do you define productivity? Akshay Nathan [01:04:46]: I think our mission is to make it possible for people to do things that they weren’t able to do before. And right now we’re thinking about it from the perspective of knowledge work. And so when I look at knowledge work, I think about people are no longer siloed by their roles. They’re no longer siloed by maybe the, background or training that they have. Like, no matter what function you’re in, you can suddenly build things. You can suddenly get access to data that you otherwise might not be able to interpret, et cetera. And then I think that extends to your personal life, where we want to give you leverage at the end of the day. Like, we want the models and the product to be able to give you leverage so that you can, create time for yourself to do the things that you love. Vibhu [01:05:25]: Does that also translate to a way to measure productivity? Like, what is new? Akshay Nathan [01:05:29]: The end is Vibhu [01:05:30]: How do you measure leverage? Akshay Nathan [01:05:31]: I think we haven’t figured this out yet. Part of the reason is it’s so diverse. Everyone has different goals, and really the true measurement is, like, their ability to achieve that goal. Did we help you or did we not? Akshay Nathan [01:05:44]: And it’s very difficult without knowing what that goal is up front and also tailoring it for every individual. Vibhu [01:05:48]: And the thumbs up and thumbs down from ChatGPT doesn’t give you anything, right? Akshay Nathan [01:05:52]: You don’t know if they’re thumbs downing the content of the answer, the vibe of it Vibhu [01:05:56]: Oh, yeah Akshay Nathan [01:05:56]: Whether or not it helped them with their goal. I think that’s difficult. But it’s something that I think we will need to figure out and the industry at large will need to figure out because, that’s how we measure success, if this is what we’re, we’re Vibhu [01:06:06]: Do you think it’s changed, productivity and how you measure it? you said there’s a lot more work that can be done, a lot more scope. has it changed? Akshay Nathan [01:06:15]: I think it was always true that what you really wanted to measure is, like, was your team, was the individual, were you personally able to hit the goal, or are you closer to hitting that, whatever your goal is, right? But I think previously we used proxies for this. So, like, code commits or Vibhu [01:06:31]: Lines of code Akshay Nathan [01:06:31]: Lines of code or whatever. Vibhu [01:06:33]: Story points. Akshay Nathan [01:06:34]: Yeah, exactly. Story points. And, like Vibhu [01:06:36]: They’re coming back, by the way. Akshay Nathan [01:06:38]: maybe. But that is for a part of the change. And, like, I think with AI now, those proxies starting to fall apart. Like, you, the number of tokens you use or the number of pull requests you make are, like, no longer, like, maybe as hypercorrelated with that, is your team able to hit the goal or are they on track to hit their goals? So I think we’ll need to come up with new, measurements. Vibhu [01:07:02]: For the managers listening, give them one thing to try. At-Bats, Motion vs. Progress, and Closing Akshay Nathan [01:07:06]: I think for me, what’s important is like at-bats. Are we as a team building the muscle to have not just quantity of at-bats, but quality? Like, are we able to go all the way from, like, generating an idea, building it out, getting the feedback, reacting to that feedback, validating or invalidating the hypothesis, going on to the next idea? Are we able to do that really efficiently? And like, that goes to like, the actual like code that’s being written or the designs that are being made or the specs that are being written, whatever, but also the culture of the team. Like, do we have the humility and, are able to like go through that process many times and stay motivated and excited throughout that? so that’s the thing that like I think is important now, especially when we’re on the frontier of this technology and like there’s so much to build, there’s so much to do. That’s probably the most important thing that we look at. Vibhu [01:07:54]: Any traps people fall into around measuring productivity with your teamwork on. I feel like there’s a lot of, okay, we added a lot of LMs. We have dashboards for this and that, but not much has changed, right? Akshay Nathan [01:08:06]: That is the trap, yes. Vibhu [01:08:09]: And the broader source of the question is for the managers and teams building, how should they approach this? Akshay Nathan [01:08:18]: I think maybe the trap is like conflating motion and progress. I think motion is much easier now than ever before because of the tooling that we have. But progress requires you to be like very prescriptive and deliberate about like what you’re trying to achieve, and it goes back to our question of measurement, right? Like you wrote we were talking about like, can we, OpenAI, like figure out how to measure productivity for our users? That’s, that’s a very hard problem because of the diversity. But like as a team, like you should have a really prescriptive and deliberate view on like what progress looks like for you and for your team. And if you don’t have that, then it’s very easy to conflate these two things. Vibhu [01:08:57]: I think at-bats is a really great thing. I’m, I’m really glad. I like the discussion between motion and progress. I think that’s a quote that we’re gonna feature on the write-up. You’ve been very generous with your time. Thank you so much and congrats on ten million. Akshay Nathan [01:09:08]: Yeah, thank you for having me. Vibhu [01:09:09]: The next one at a hundred in two months. Two weeks. Thank you. This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.latent.space/subscribe
The podcast currently has 249 episodes available.

1,091 Listeners

303 Listeners

338 Listeners

225 Listeners

205 Listeners

204 Listeners

317 Listeners

99 Listeners

582 Listeners

503 Listeners

142 Listeners

222 Listeners

684 Listeners

456 Listeners

30 Listeners