AI Post Transformers

AgenticQwen and Small Industrial Tool Agents


Listen Later

This episode explores AgenticQwen, a system for training small open-weight language models to handle industrial-scale tool use through repeated reinforcement learning and synthetic data generation. It explains the paper’s central idea of dual data flywheels: one that turns reasoning failures into harder verifiable tasks, and another that expands simple agent workflows into branching, tool-using trajectories with recovery steps and user interaction. The discussion contrasts imitation from synthetic data with trajectory-level reinforcement learning, arguing that real agent competence depends on rewarding decisions like tool choice, clarification, and error recovery rather than just polished final answers. Listeners would find it interesting for its grounded look at whether small models can become cheap, fast, and genuinely useful agents for high-volume real-world work without relying on massive frontier systems.
Sources:
1. AgenticQwen: Training Small Agentic Language Models with Dual Data Flywheels for Industrial-Scale Tool Use — Yuanjie Lyu, Chengyu Wang, Haonan Zheng, Yuanhao Yue, Junbing Yan, Ming Wang, Jun Huang, 2026
http://arxiv.org/abs/2604.21590
2. Training Language Models to Follow Instructions with Human Feedback — Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin and others, 2022
https://scholar.google.com/scholar?q=Training+Language+Models+to+Follow+Instructions+with+Human+Feedback
3. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning — DeepSeek-AI, 2025
https://scholar.google.com/scholar?q=DeepSeek-R1%3A+Incentivizing+Reasoning+Capability+in+LLMs+via+Reinforcement+Learning
4. Agent Lightning: Train ANY AI Agents with Reinforcement Learning — Xufang Luo, Yuge Zhang, Zhiyuan He, Zilong Wang, Siyun Zhao, Dongsheng Li, Luna K. Qiu, Yuqing Yang and others, 2025
https://scholar.google.com/scholar?q=Agent+Lightning%3A+Train+ANY+AI+Agents+with+Reinforcement+Learning
5. CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use — Zhen Zhang, Kaiqiang Song, Xun Wang, Yebowen Hu, Weixiang Yan, Chenyang Zhao, Henry Peng Zou and others, 2026
https://scholar.google.com/scholar?q=CM2%3A+Reinforcement+Learning+with+Checklist+Rewards+for+Multi-Turn+and+Multi-Step+Agentic+Tool+Use
6. Self-Instruct: Aligning Language Models with Self-Generated Instructions — Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, Hannaneh Hajishirzi, 2022
https://scholar.google.com/scholar?q=Self-Instruct%3A+Aligning+Language+Models+with+Self-Generated+Instructions
7. Orca: Progressive Learning from Complex Explanation Traces of GPT-4 — Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar, Sahaj Agarwal, Hamid Palangi, Ahmed Awadallah, 2023
https://scholar.google.com/scholar?q=Orca%3A+Progressive+Learning+from+Complex+Explanation+Traces+of+GPT-4
8. Textbooks Are All You Need — Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio Cesar Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi and others, 2023
https://scholar.google.com/scholar?q=Textbooks+Are+All+You+Need
9. TinyStories: How Small Can Language Models Be and Still Speak Coherent English? — Ronen Eldan, Yuanzhi Li, 2023
https://scholar.google.com/scholar?q=TinyStories%3A+How+Small+Can+Language+Models+Be+and+Still+Speak+Coherent+English%3F
10. A Survey of Behavior Trees in Robotics and AI — Matteo Iovino, Edvards Scukins, Jonathan Styrud, Petter Ogren, Christian Smith, 2022
https://scholar.google.com/scholar?q=A+Survey+of+Behavior+Trees+in+Robotics+and+AI
11. Robot Behavior-Tree-Based Task Generation with Large Language Models — Yue Cao, C. S. George Lee, 2023
https://scholar.google.com/scholar?q=Robot+Behavior-Tree-Based+Task+Generation+with+Large+Language+Models
12. Behavior Trees Enable Structured Programming of Language Model Agents — Richard Kelley, 2024
https://scholar.google.com/scholar?q=Behavior+Trees+Enable+Structured+Programming+of+Language+Model+Agents
13. Behavior Tree Generation and Adaptation for a Social Robot Control with LLMs — authors from the 2025 Robotics and Autonomous Systems paper, 2025
https://scholar.google.com/scholar?q=Behavior+Tree+Generation+and+Adaptation+for+a+Social+Robot+Control+with+LLMs
14. Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewards — Yuan-Jay Lü, Chengyu Wang, Lei Shen, Jun Huang, Tong Xu, 2026
https://scholar.google.com/scholar?q=Mock+Worlds%2C+Real+Skills%3A+Building+Small+Agentic+Language+Models+with+Synthetic+Tasks%2C+Simulated+Environments%2C+and+Rubric-Based+Rewards
15. τ^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment — Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Karthik Narasimhan, 2025
https://scholar.google.com/scholar?q=%CF%84%5E2-Bench%3A+Evaluating+Conversational+Agents+in+a+Dual-Control+Environment
16. Procedural Environment Generation for Tool-Use Agents — Michael Sullivan, Mareike Hartmann, Alexander Koller, 2025
https://scholar.google.com/scholar?q=Procedural+Environment+Generation+for+Tool-Use+Agents
17. ASTRA-bench: Evaluating Tool-Use Agent Reasoning and Action Planning with Personal User Context — approx. ASTRA-bench authors unknown from snippet, recent
https://scholar.google.com/scholar?q=ASTRA-bench%3A+Evaluating+Tool-Use+Agent+Reasoning+and+Action+Planning+with+Personal+User+Context
18. m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks — approx. m&m's benchmark authors unknown from snippet, recent
https://scholar.google.com/scholar?q=m%26m%27s%3A+A+Benchmark+to+Evaluate+Tool-Use+for+multi-step+multi-modal+Tasks
19. EvilGenie: A Reward Hacking Benchmark — approx. EvilGenie authors unknown from snippet, recent
https://scholar.google.com/scholar?q=EvilGenie%3A+A+Reward+Hacking+Benchmark
20. Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking — approx. reward auditing authors unknown from snippet, recent
https://scholar.google.com/scholar?q=Adversarial+Reward+Auditing+for+Active+Detection+and+Mitigation+of+Reward+Hacking
21. AI Post Transformers: Benchmarking Test-Time Scaling for General LLM Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-benchmarking-test-time-scaling-for-gener-8f14f9.mp3
22. AI Post Transformers: Experience-Based Learning Beyond Human Data — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-experience-based-learning-beyond-human-d-b0caa4.mp3
23. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
24. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3
25. AI Post Transformers: Simple Self-Distillation for Better Code Generation — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-02-simple-self-distillation-for-better-code-cc88e0.mp3
Interactive Visualization: AgenticQwen and Small Industrial Tool Agents
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof