This episode explores Apple’s paper on whether code models can improve through an extremely simple form of self-distillation: fine-tuning on their own sampled code outputs without using a stronger teacher, execution feedback, verifiers, or reinforcement learning. It situates that idea within the broader history of knowledge distillation and post-training, comparing it to earlier work like Hinton’s distillation, sequence-level distillation, Born Again Networks, Noisy Student, and newer on-policy language model distillation. The discussion focuses on why code generation is a particularly revealing testbed, since benchmarks like pass@1 and pass@k make it easier to tell whether self-distillation is uncovering latent capability or just repackaging errors. A listener would find it interesting because the paper challenges a core assumption in modern model improvement: that meaningful gains require expensive external supervision rather than a surprisingly cheap training loop around the model itself.
Sources:
1. Embarrassingly Simple Self-Distillation Improves Code Generation — Ruixiang Zhang, Richard He Bai, Huangjie Zheng, Navdeep Jaitly, Ronan Collobert, Yizhe Zhang, 2026
http://arxiv.org/abs/2604.01193
2. Distilling the Knowledge in a Neural Network — Geoffrey Hinton, Oriol Vinyals, Jeff Dean, 2015
https://scholar.google.com/scholar?q=Distilling+the+Knowledge+in+a+Neural+Network
3. Sequence-Level Knowledge Distillation — Yoon Kim, Alexander M. Rush, 2016
https://scholar.google.com/scholar?q=Sequence-Level+Knowledge+Distillation
4. Born Again Neural Networks — Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, Anima Anandkumar, 2018
https://scholar.google.com/scholar?q=Born+Again+Neural+Networks
5. Self-training with Noisy Student improves ImageNet classification — Qizhe Xie, Minh-Thang Luong, Eduard Hovy, Quoc V. Le, 2020
https://scholar.google.com/scholar?q=Self-training+with+Noisy+Student+improves+ImageNet+classification
6. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes — Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, Olivier Bachem, 2024
https://scholar.google.com/scholar?q=On-Policy+Distillation+of+Language+Models%3A+Learning+from+Self-Generated+Mistakes
7. Evaluating Large Language Models Trained on Code — Mark Chen, Jerry Tworek, Heewoo Jun, et al., 2021
https://scholar.google.com/scholar?q=Evaluating+Large+Language+Models+Trained+on+Code
8. Program Synthesis with Large Language Models — Jacob Austin, Augustus Odena, Maxwell Nye, et al., 2021
https://scholar.google.com/scholar?q=Program+Synthesis+with+Large+Language+Models
9. Measuring Coding Challenge Competence With APPS — Dan Hendrycks, Collin Burns, Steven Basart, et al., 2021
https://scholar.google.com/scholar?q=Measuring+Coding+Challenge+Competence+With+APPS
10. Self-Consistency Improves Chain of Thought Reasoning in Language Models — Xuezhi Wang, Jason Wei, Dale Schuurmans, et al., 2022
https://scholar.google.com/scholar?q=Self-Consistency+Improves+Chain+of+Thought+Reasoning+in+Language+Models
11. DeepSeek-R1 — DeepSeek-AI, 2025
https://scholar.google.com/scholar?q=DeepSeek-R1
12. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code — Prasenjit Jain, et al., 2024
https://scholar.google.com/scholar?q=LiveCodeBench%3A+Holistic+and+Contamination+Free+Evaluation+of+Large+Language+Models+for+Code
13. SelfCodeAlign: Self-Alignment for Code Generation — Yuxiang Wei, Federico Cassano, Jiawei Liu, Yifeng Ding, Naman Jain, Zachary Mueller, Harm de Vries, Leandro von Werra, Arjun Guha, Lingming Zhang, 2024
https://scholar.google.com/scholar?q=SelfCodeAlign%3A+Self-Alignment+for+Code+Generation
14. Iterative Self-Training for Code Generation via Reinforced Re-Ranking — Nikita Sorokin, Ivan Sedykh, Valentin Malykh, 2025
https://scholar.google.com/scholar?q=Iterative+Self-Training+for+Code+Generation+via+Reinforced+Re-Ranking
15. On the Role of Temperature Sampling in Test-Time Scaling — Yuheng Wu, Azalia Mirhoseini, Thierry Tambe, 2025
https://scholar.google.com/scholar?q=On+the+Role+of+Temperature+Sampling+in+Test-Time+Scaling
16. OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement — Tianyu Zheng, Ge Zhang, Tianhao Shen, Xueling Liu, Bill Yuchen Lin, Jie Fu, Wenhu Chen, Xiang Yue, 2024
https://scholar.google.com/scholar?q=OpenCodeInterpreter%3A+Integrating+Code+Generation+with+Execution+and+Refinement
17. GenX: Mastering Code and Test Generation with Execution Feedback — Nan Wang, Yafei Liu, Chen Chen, Haonan Lu, 2024
https://scholar.google.com/scholar?q=GenX%3A+Mastering+Code+and+Test+Generation+with+Execution+Feedback
18. InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback — John Yang, Akshara Prabhakar, Karthik Narasimhan, Shunyu Yao, 2023
https://scholar.google.com/scholar?q=InterCode%3A+Standardizing+and+Benchmarking+Interactive+Coding+with+Execution+Feedback
19. AI Post Transformers: Evolving Language Models Without Labels: EVOL-RL — Hal Turing & Dr. Ada Shannon, Fri,
https://podcast.do-not-panic.com/episodes/evolving-language-models-without-labels-evol-rl/
20. AI Post Transformers: Lp-Reg: Low-Probability Tokens Sustain RL Exploration — Hal Turing & Dr. Ada Shannon, Sun,
https://podcast.do-not-panic.com/episodes/lp-reg-low-probability-tokens-sustain-rl-exploration/
21. AI Post Transformers: NeurIPS 2025: SeRL: Self-Play Reinforcement Learning for Large Language Models with Limited Data — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/neurips-2025-serl-self-play-reinforcement-learning-for-large-language-models-wit/
22. AI Post Transformers: LLM Benchmark Robustness to Linguistic Variation — Hal Turing & Dr. Ada Shannon, Tue,
https://podcast.do-not-panic.com/episodes/llm-benchmark-robustness-to-linguistic-variation/
Interactive Visualization: Simple Self-Distillation for Better Code Generation