This episode explores OpenSkill, a framework for LLM agents that tries to improve behavior after deployment by building durable, reusable skills from public evidence rather than retraining model weights. It explains how the paper separates ordinary tool use from open-world self-evolution, arguing that the key challenge is not just acting with browsers and code, but turning documentation, repositories, papers, and tutorials into explicit procedures and verification checks. The discussion focuses on the paper’s central claim that agents can create their own proxy tests through grounded verification anchors without leaking hidden benchmark answers, and compares that approach with earlier systems like Reflexion, Voyager, ExpeL, AutoSkill, and Memento-Skills. Listeners would find it interesting because it gets at a practical industry problem: whether agents can stay useful as APIs, websites, and workflows change, or whether the verifier remains the real bottleneck to genuine self-improvement.
Sources:
1. OpenSkill: Open-World Self-Evolution for LLM Agents — Zhiling Yan, Dingjie Song, Hanrong Zhang, Wei Liang, Yuxuan Zhang, Yutong Dai, Lifang He, Philip S. Yu, Ran Xu, Xiang Li, Lichao Sun, 2026
http://arxiv.org/abs/2606.06741
2. Reflexion: Language Agents with Verbal Reinforcement Learning — Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, Shunyu Yao, 2023
https://arxiv.org/abs/2303.11366
3. Voyager: An Open-Ended Embodied Agent with Large Language Models — Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, Anima Anandkumar, 2023
https://arxiv.org/abs/2305.16291
4. ExpeL: LLM Agents Are Experiential Learners — Andrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin, Yong-Jin Liu, Gao Huang, 2023
https://arxiv.org/abs/2308.10144
5. Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration — Qifan Zhang, Dongyang Ma, Tianqing Fang, Jia Li, Jing Tang, Nuo Chen, Haitao Mi, Yan Wang, 2026
https://arxiv.org/abs/2604.18131
6. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Saurav Kadavath, Amanda Askell, Ethan Perez, Jared Kaplan, Dario Amodei, Tom Brown and collaborators, 2022
https://arxiv.org/abs/2212.08073
7. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback — Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, Sushant Prakash, 2023
https://arxiv.org/abs/2309.00267
8. Self-Rewarding Language Models — Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li, Sainbayar Sukhbaatar, Jing Xu, Jason Weston, 2024
https://arxiv.org/abs/2401.10020
9. MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models — Longhui Yu, Weisen Jiang, Han Shi, Jincheng Yu, Zhengying Liu, Yu Zhang, James T. Kwok, Zhenguo Li, Adrian Weller, Weiyang Liu, 2023
https://arxiv.org/abs/2309.12284
10. SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks — Xiangyi Li et al., 2026
https://scholar.google.com/scholar?q=SkillsBench%3A+Benchmarking+How+Well+Agent+Skills+Work+Across+Diverse+Tasks
11. AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution — Yutao Yang et al., 2026
https://scholar.google.com/scholar?q=AutoSkill%3A+Experience-Driven+Lifelong+Learning+via+Skill+Self-Evolution
12. Memento-Skills: Let Agents Design Agents — Huichi Zhou et al., 2026
https://scholar.google.com/scholar?q=Memento-Skills%3A+Let+Agents+Design+Agents
13. SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces — Chang Jin et al., 2026
https://scholar.google.com/scholar?q=SkillSafetyBench%3A+Evaluating+Agent+Safety+under+Skill-Facing+Attack+Surfaces
14. EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction — Siyu Yuan et al., 2024
https://scholar.google.com/scholar?q=EASYTOOL%3A+Enhancing+LLM-based+Agents+with+Concise+Tool+Instruction
15. AutoRefine: From Trajectories to Reusable Expertise for Continual LLM Agent Refinement — Libin Qiu et al., 2026
https://scholar.google.com/scholar?q=AutoRefine%3A+From+Trajectories+to+Reusable+Expertise+for+Continual+LLM+Agent+Refinement
16. SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents — Jiaye Lin et al., 2025
https://scholar.google.com/scholar?q=SE-Agent%3A+Self-Evolution+Trajectory+Optimization+in+Multi-Step+Reasoning+with+LLM-Based+Agents
17. When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs — Fangyi Yu, 2025
https://scholar.google.com/scholar?q=When+AIs+Judge+AIs%3A+The+Rise+of+Agent-as-a-Judge+Evaluation+for+LLMs
18. Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments — Yuran Li et al., 2025
https://scholar.google.com/scholar?q=Leveraging+LLMs+as+Meta-Judges%3A+A+Multi-Agent+Framework+for+Evaluating+LLM+Judgments
19. AI Post Transformers: The Endless Gym: Training Terminal Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/the-endless-gym-training-terminal-agents/
20. AI Post Transformers: When AI Builds Itself and Recursive Self-Improvement — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-05-when-ai-builds-itself-and-recursive-self-8bbf9e.mp3
21. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3
22. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3
23. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
Interactive Visualization: OpenSkill for Open-World Self-Evolution in LLM Agents