This episode explores Marina Favaro et al.’s 2026 paper on whether AI is beginning to accelerate frontier AI development enough to hint at recursive self-improvement, laying out the ladder from chatbots and coding agents to systems that materially help build their own successors. It distinguishes progress on long-horizon, fixed-goal engineering tasks from the harder problem of genuine research judgment, using public evidence such as METR, CORE-Bench, and RE-Bench to argue that AI is clearly getting better at sustained execution but has not yet shown strong scientific taste or reliable autonomy. It then digs into Anthropic’s internal evidence, including claims that by May 2026 Claude was responsible for over 80% of merged production code, open-ended coding-task success had risen sharply, and fixed-goal research engineering tasks improved from roughly 3x to 52x over a year, while the speakers repeatedly stress that code volume and self-reported productivity overstate true impact. Listeners would find it interesting because the discussion ties concrete benchmark results and lab productivity numbers to the bigger question of whether faster AI development also compresses the timelines for safety, governance, and the arrival of more capable successor systems.
Sources:
1. When AI Builds Itself and Recursive Self-Improvement
https://www.anthropic.com/institute/recursive-self-improvement
2. Measuring AI Ability to Complete Long Tasks — Thomas Kwa, Ben West, Joel Becker, Amy Deng, et al., 2025
https://scholar.google.com/scholar?q=Measuring+AI+Ability+to+Complete+Long+Tasks
3. RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts — Hjalmar Wijk, Tao Roa Lin, Joel Becker, Sami Jawhar, et al., 2025
https://scholar.google.com/scholar?q=RE-Bench%3A+Evaluating+Frontier+AI+R%26D+Capabilities+of+Language+Model+Agents+against+Human+Experts
4. CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark — Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, Arvind Narayanan, 2024
https://scholar.google.com/scholar?q=CORE-Bench%3A+Fostering+the+Credibility+of+Published+Research+Through+a+Computational+Reproducibility+Agent+Benchmark
5. Automated Weak-to-Strong Researcher — Jiaxin Wen, Liang Qiu, Joe Benton, Jan Hendrik Kirchner, Jan Leike, 2026
https://scholar.google.com/scholar?q=Automated+Weak-to-Strong+Researcher
6. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, et al., 2025
https://scholar.google.com/scholar?q=PaperBench%3A+Evaluating+AI%27s+Ability+to+Replicate+AI+Research
7. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — Joel Becker, Nate Rush, Elizabeth Barnes, David Rein, 2025
https://scholar.google.com/scholar?q=Measuring+the+Impact+of+Early-2025+AI+on+Experienced+Open-Source+Developer+Productivity
8. SWE-bench Goes Live! — Linghao Zhang et al., 2025
https://scholar.google.com/scholar?q=SWE-bench+Goes+Live%21
9. Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving — Daoguang Zan et al., 2025
https://scholar.google.com/scholar?q=Multi-SWE-bench%3A+A+Multilingual+Benchmark+for+Issue+Resolving
10. SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? — John Yang et al., 2024
https://scholar.google.com/scholar?q=SWE-bench+Multimodal%3A+Do+AI+Systems+Generalize+to+Visual+Software+Domains%3F
11. Can AI Conduct Autonomous Scientific Research? Case Studies on Two Real-World Tasks — S. Agrawal et al., 2026
https://scholar.google.com/scholar?q=Can+AI+Conduct+Autonomous+Scientific+Research%3F+Case+Studies+on+Two+Real-World+Tasks
12. Collapse of Self-trained Language Models — David Herel and Tomas Mikolov, 2024
https://scholar.google.com/scholar?q=Collapse+of+Self-trained+Language+Models
13. Exploring Automation Bias in Human-AI Collaboration: A Review and Implications for Explainable AI — G. Romeo and D. Conti, 2025
https://scholar.google.com/scholar?q=Exploring+Automation+Bias+in+Human-AI+Collaboration%3A+A+Review+and+Implications+for+Explainable+AI
14. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
15. AI Post Transformers: Trajectory Summaries for Long-Horizon Coding Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-24-trajectory-summaries-for-long-horizon-co-0194be.mp3
16. AI Post Transformers: TMAS: Scaling Test-Time Compute with Multi-Agent Synergy — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-14-tmas-scaling-test-time-compute-with-mult-3abe7a.mp3
17. AI Post Transformers: Air Force One, Jensen Huang, and Anthropic's 2028 Memo — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-2028-scenarios-for-global-ai-leadership-d0ec29.mp3
Interactive Visualization: When AI Builds Itself and Recursive Self-Improvement