AI Post Transformers

Self-Improving Pretraining With Post-Trained Models


Listen Later

This episode explores the idea of “self-improving pretraining,” where already post-trained models are used to shape the pretraining of new models rather than waiting to add safety, reasoning, and factuality later. It explains how the approach rewrites training continuations, uses stronger models as judges, and compares original corpus text, teacher-generated suffixes, and learner rollouts to push model preferences upstream into training. The discussion also situates the paper against earlier work like Constitutional AI, STaR, and Quiet-STaR, while debating whether this is a genuine shift in training philosophy or mainly a more aggressive form of distilling a stronger model’s preferences. Listeners would find it interesting because it gets at a central question in modern AI: whether better behavior can be built into a model’s foundations instead of patched on after the fact.
Sources:
1. Self-Improving Pretraining: using post-trained models to pretrain better models — Ellen Xiaoqing Tan, Jack Lanchantin, Shehzaad Dhuliawala, Danwei Li, Thao Nguyen, Jing Xu, Ping Yu, Ilia Kulikov, Sainbayar Sukhbaatar, Jason Weston, Xian Li, Olga Golovneva, 2026
http://arxiv.org/abs/2601.21343
2. Constitutional AI: Harmlessness from AI Feedback — Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, et al., 2022
https://scholar.google.com/scholar?q=Constitutional+AI%3A+Harmlessness+from+AI+Feedback
3. STaR: Self-Taught Reasoner Bootstrapping Reasoning With Reasoning — Eric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. Goodman, and Percy Liang, 2022
https://scholar.google.com/scholar?q=STaR%3A+Self-Taught+Reasoner+Bootstrapping+Reasoning+With+Reasoning
4. Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking — Eric Zelikman, Yuhuai Wu, Noah D. Goodman, and Jakob Foerster, 2024
https://scholar.google.com/scholar?q=Quiet-STaR%3A+Language+Models+Can+Teach+Themselves+to+Think+Before+Speaking
5. Self-Rewarding Language Models — Natasha Shumailov, John Adler, Ming-Wei Chang, Sharan Narang, and Yi Tay, 2024
https://scholar.google.com/scholar?q=Self-Rewarding+Language+Models
6. Confronting Reward Model Overoptimization with Constrained RLHF — Nathan Lambert, Louis Castricato, Leandro von Werra, and Alex Havrilla, 2024
https://scholar.google.com/scholar?q=Confronting+Reward+Model+Overoptimization+with+Constrained+RLHF
7. On the Diversity of Synthetic Data and its Impact on Training Large Language Models — Hao Chen, Abdul Waheed, Xiang Li, Yidong Wang, Jindong Wang, Bhiksha Raj, Marah I. Abdin, 2024
https://scholar.google.com/scholar?q=On+the+Diversity+of+Synthetic+Data+and+its+Impact+on+Training+Large+Language+Models
8. In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners — Jaehoon Kim, Kwangwook Seo, Dongha Lee, 2025
https://scholar.google.com/scholar?q=In+Their+Own+Words%3A+Reasoning+Traces+Tailored+for+Small+Models+Make+Them+Better+Reasoners
9. Rethinking Rubric Generation for Improving LLM Judge and Reward Modeling for Open-ended Tasks — William F. Shen, Xinchi Qiu, Chenxi Whitehouse, Lisa Alazraki, Shashwat Goel, Francesco Barbieri, Timon Willi, Akhil Mathur, Ilias Leontiadis, 2026
https://scholar.google.com/scholar?q=Rethinking+Rubric+Generation+for+Improving+LLM+Judge+and+Reward+Modeling+for+Open-ended+Tasks
10. Ask a Strong LLM Judge when Your Reward Model is Uncertain — Zhenghao Xu, Qin Lu, Qingru Zhang, Liang Qiu, Ilgee Hong, Changlong Yu, Wenlin Yao, Yao Liu, Haoming Jiang, Lihong Li, Hyokun Yun, Tuo Zhao, 2025
https://scholar.google.com/scholar?q=Ask+a+Strong+LLM+Judge+when+Your+Reward+Model+is+Uncertain
11. FOL-Traces: Verified First-Order Logic Reasoning Traces at Scale — Isabelle Lee, Sarah Liaw, Dani Yogatama, 2025
https://scholar.google.com/scholar?q=FOL-Traces%3A+Verified+First-Order+Logic+Reasoning+Traces+at+Scale
12. AI Post Transformers: ASI-Evolve for Data, Architectures, and RL — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-asi-evolve-for-data-architectures-and-rl-197b2b.mp3
13. AI Post Transformers: Distilling Multi-Agent Reasoning into a Single LLM — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-distilling-multi-agent-reasoning-into-a-143263.mp3
14. AI Post Transformers: Learning to Reason with 13 Parameters — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-learning-to-reason-with-13-parameters-54c87f.mp3
15. AI Post Transformers: MASA: Meta-Awareness via Self-Alignment Reinforcement Learning — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/masa-meta-awareness-via-self-alignment-reinforcement-learning/
16. AI Post Transformers: Kimi K2.5 and Visual Agent Swarms — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-24-kimi-k25-and-visual-agent-swarms-7d04d7.mp3
Interactive Visualization: Self-Improving Pretraining With Post-Trained Models
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof