AI Post Transformers

PaperBench: Can AI Replicate AI Research?


Listen Later

This episode explores PaperBench, a benchmark designed to test whether frontier AI agents can independently replicate the empirical work of recent machine learning papers from scratch rather than merely explain them. It breaks down what agentic AI actually entails in this setting: reading papers, writing code, choosing baselines, reconstructing missing details, running experiments, debugging failures, and judging whether reproduced results match the original claims. The discussion compares PaperBench with other evaluation ladders such as CORE-Bench, MLE-bench, RE-Bench, and JudgeEval, while also debating whether controlled scratch replication should be viewed as advanced engineering or a meaningful proxy for real research practice. Listeners get a clear look at why this matters for both AI capability measurement and safety, especially given PaperBench’s carefully curated design of 20 ICML 2024 papers, 12 topics, and more than 8,000 graded tasks.
Sources:
1. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan, Leon Maksin, Rachel Dias, Evan Mays, Benjamin Kinsella, Wyatt Thompson, Johannes Heidecke, Amelia Glaese, Tejal Patwardhan, 2025
http://arxiv.org/abs/2504.01848
2. PaperBench: Evaluating AI's Ability to Replicate AI Research — Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, et al., 2025
https://scholar.google.com/scholar?q=PaperBench%3A+Evaluating+AI%27s+Ability+to+Replicate+AI+Research
3. RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts — Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, et al., 2024
https://scholar.google.com/scholar?q=RE-Bench%3A+Evaluating+frontier+AI+R%26D+capabilities+of+language+model+agents+against+human+experts
4. CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark — Zachary S. Siegel, Sayash Kapoor, Nitya Nagdir, Benedikt Stroebl, Arvind Narayanan, 2024
https://scholar.google.com/scholar?q=CORE-Bench%3A+Fostering+the+Credibility+of+Published+Research+Through+a+Computational+Reproducibility+Agent+Benchmark
5. MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation — Qian Huang, Jian Vora, Percy Liang, Jure Leskovec, 2023
https://scholar.google.com/scholar?q=MLAgentBench%3A+Evaluating+Language+Agents+on+Machine+Learning+Experimentation
6. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering — Jun Shern Chan et al., 2024
https://scholar.google.com/scholar?q=MLE-bench%3A+Evaluating+Machine+Learning+Agents+on+Machine+Learning+Engineering
7. EXP-Bench: Can AI Conduct AI Research Experiments? — Patrick Tser Jern Kon et al., 2025
https://scholar.google.com/scholar?q=EXP-Bench%3A+Can+AI+Conduct+AI+Research+Experiments%3F
8. MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research — Hui Chen, Miao Xiong, Yujie Lu, Wei Han, Ailin Deng, Yufei He, Jiaying Wu, Yibo Li, Yue Liu, Bryan Hooi, 2025
https://scholar.google.com/scholar?q=MLR-Bench%3A+Evaluating+AI+Agents+on+Open-Ended+Machine+Learning+Research
9. ReplicationBench: Can AI Agents Replicate Astrophysics Research Papers? — Christine Ye et al., 2025
https://scholar.google.com/scholar?q=ReplicationBench%3A+Can+AI+Agents+Replicate+Astrophysics+Research+Papers%3F
10. Can Large Language Models Be an Alternative to Human Evaluations? — Cheng-Han Chiang, Hung-yi Lee, 2023
https://scholar.google.com/scholar?q=Can+Large+Language+Models+Be+an+Alternative+to+Human+Evaluations%3F
11. RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following — Tianjun Pan et al., 2026
https://arxiv.org/abs/2603.25133
12. JudgeBench: A Benchmark for Evaluating LLM-based Judges — Sijun Tan et al., 2024
https://arxiv.org/abs/2410.12784
13. When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation — Abeer Badawi et al., 2025
https://arxiv.org/abs/2510.19032
14. A Dataset For Computational Reproducibility — Lazaro Costa, Susana Barbosa, Jacome Cunha, 2025
https://arxiv.org/abs/2504.08684
15. SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers — Yanzheng Xiang et al., 2025
https://arxiv.org/abs/2504.00255
16. OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding — Deming Ding et al., 2026
https://arxiv.org/abs/2601.10343
17. ContextBench: A Benchmark for Context Retrieval in Coding Agents — Han Li et al., 2026
https://arxiv.org/abs/2602.05892
18. AI Post Transformers: When AI Builds Itself and Recursive Self-Improvement — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-when-ai-builds-itself-and-recursive-self-8bbf9e.mp3
19. AI Post Transformers: When LLM Judges Become Coin Flips — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-when-llm-judges-become-coin-flips-8b43ef.mp3
20. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
21. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3
Interactive Visualization: PaperBench: Can AI Replicate AI Research?
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof