This episode explores a 2026 paper on defending language models against unauthorized distillation by rewriting chain-of-thought traces before they are returned through an API. It explains the core idea of making reasoning outputs remain useful and correct for human users while becoming less effective as training data for a smaller model trying to copy the teacher, and it connects that strategy to data poisoning and watermarking. The discussion focuses on two defense families, especially LLM-based trace rewriting, and highlights reported results showing strong student degradation on reasoning-heavy tasks like MATH while often preserving or even improving teacher performance. It also digs into the paper’s main ambiguity: whether the defense truly poisons the student’s learning signal, or whether a stronger rewrite model is simply producing cleaner, differently structured reasoning that smaller distilled models fail to absorb well.
Sources:
1. Protecting Language Models Against Unauthorized Distillation through Trace Rewriting — Xinhang Ma, William Yeoh, Ning Zhang, Yevgeniy Vorobeychik, 2026
http://arxiv.org/abs/2602.15143
2. Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? — Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu, Xuming Hu, Lijie Wen, Irwin King, Philip S. Yu, 2025
http://arxiv.org/abs/2502.11598
3. A Survey of Deep Neural Network Watermarking Techniques — Yusuke Li, Huili Wang, Mauro Barni, 2021
https://scholar.google.com/scholar?q=A+Survey+of+Deep+Neural+Network+Watermarking+Techniques
4. Distillation-Resistant Watermarking for Model Protection in NLP — Xuandong Zhao, Lei Li, Yu-Xiang Wang, 2022
https://scholar.google.com/scholar?q=Distillation-Resistant+Watermarking+for+Model+Protection+in+NLP
5. Protecting Language Generation Models via Invisible Watermarking — Xuandong Zhao, Yu-Xiang Wang, Lei Li, 2023
https://scholar.google.com/scholar?q=Protecting+Language+Generation+Models+via+Invisible+Watermarking
6. Scalable Watermarking for Identifying Large Language Model Outputs — Sumanth Dathathri, Abigail See, S. Ghaisas and colleagues, 2024
https://scholar.google.com/scholar?q=Scalable+Watermarking+for+Identifying+Large+Language+Model+Outputs
7. Poisoning Attacks against Support Vector Machines — Battista Biggio, Blaine Nelson, Pavel Laskov, 2012
https://scholar.google.com/scholar?q=Poisoning+Attacks+against+Support+Vector+Machines
8. Dataset Security for Machine Learning: Data Poisoning, Backdoor Attacks, and Defenses — Micah Goldblum and colleagues, 2020
https://scholar.google.com/scholar?q=Dataset+Security+for+Machine+Learning%3A+Data+Poisoning%2C+Backdoor+Attacks%2C+and+Defenses
9. Online Data Poisoning Attacks — Xuezhou Zhang, Xiaojin Zhu, Laurent Lessard, 2020
https://scholar.google.com/scholar?q=Online+Data+Poisoning+Attacks
10. Poisoning Language Models During Instruction Tuning — Alexander Wan, Eric Wallace, Sheng Shen, Dan Klein, 2023
https://scholar.google.com/scholar?q=Poisoning+Language+Models+During+Instruction+Tuning
11. Adversarial Training Methods for Semi-Supervised Text Classification — Takeru Miyato, Andrew M. Dai, Ian Goodfellow, 2017
https://scholar.google.com/scholar?q=Adversarial+Training+Methods+for+Semi-Supervised+Text+Classification
12. HotFlip: White-Box Adversarial Examples for Text Classification — Javid Ebrahimi, Anyi Rao, Daniel Lowd, Dejing Dou, 2018
https://scholar.google.com/scholar?q=HotFlip%3A+White-Box+Adversarial+Examples+for+Text+Classification
13. Universal Adversarial Triggers for Attacking and Analyzing NLP — Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, Sameer Singh, 2019
https://scholar.google.com/scholar?q=Universal+Adversarial+Triggers+for+Attacking+and+Analyzing+NLP
14. Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework — Lifan Yuan, Yichi Zhang, Yangyi Chen, Wei Wei, 2023
https://scholar.google.com/scholar?q=Bridge+the+Gap+Between+CV+and+NLP%21+A+Gradient-based+Textual+Adversarial+Attack+Framework
15. Antidistillation Sampling — Yash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu, Avi Schwarzschild, Alexander Robey, Marc Finzi, J. Zico Kolter, 2025
https://scholar.google.com/scholar?q=Antidistillation+Sampling
16. DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation — Pingzhi Li, Zhen Tan, Mohan Zhang, Huaizhi Qu, Huan Liu, Tianlong Chen, 2025
https://scholar.google.com/scholar?q=DOGe%3A+Defensive+Output+Generation+for+LLM+Protection+Against+Knowledge+Distillation
17. Information-Preserving Reformulation of Reasoning Traces for Antidistillation — Jiayu Ding, Lei Cui, Li Dong, Nanning Zheng, Furu Wei, 2025
https://scholar.google.com/scholar?q=Information-Preserving+Reformulation+of+Reasoning+Traces+for+Antidistillation
18. Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation? — Leyi Pan, Aiwei Liu, Shiyu Huang, Yijian Lu, Xuming Hu, Lijie Wen, Irwin King, Philip S. Yu, 2025
https://scholar.google.com/scholar?q=Can+LLM+Watermarks+Robustly+Prevent+Unauthorized+Knowledge+Distillation%3F
19. Unified Attacks to Large Language Model Watermarks: Spoofing and Scrubbing in Unauthorized Knowledge Distillation — Xin Yi, Yue Li, Shunfan Zheng, Linlin Wang, Xiaoling Wang, Liang He, 2025
https://scholar.google.com/scholar?q=Unified+Attacks+to+Large+Language+Model+Watermarks%3A+Spoofing+and+Scrubbing+in+Unauthorized+Knowledge+Distillation
20. CATER: Intellectual Property Protection on Text Generation APIs via Conditional Watermarks — Xuanli He, Qiongkai Xu, Yi Zeng, Lingjuan Lyu, Fangzhao Wu, Jiwei Li, Ruoxi Jia, 2022
https://scholar.google.com/scholar?q=CATER%3A+Intellectual+Property+Protection+on+Text+Generation+APIs+via+Conditional+Watermarks
21. DistillGuard: Evaluating Defenses Against LLM Knowledge Distillation — Bo Jiang, 2026
https://scholar.google.com/scholar?q=DistillGuard%3A+Evaluating+Defenses+Against+LLM+Knowledge+Distillation
22. Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational Agents — Hyungjoo Chae, Yongho Song, Kai Tzu-iunn Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee, Dongyeop Kang, Jinyoung Yeo, 2023
https://scholar.google.com/scholar?q=Dialogue+Chain-of-Thought+Distillation+for+Commonsense-aware+Conversational+Agents
23. Towards Faithful Multi-step Reasoning through Fine-Grained Causal-aware Attribution Reasoning Distillation — Zheng Chu, Jingchang Chen, Zhongjie Wang, Guo Tang, Qianglong Chen, Ming Liu, Bing Qin, 2025
https://scholar.google.com/scholar?q=Towards+Faithful+Multi-step+Reasoning+through+Fine-Grained+Causal-aware+Attribution+Reasoning+Distillation
24. Large Language Model Watermark Stealing With Mixed Integer Programming — Zhaoxi Zhang, Xiaomei Zhang, Yanjun Zhang, Leo Yu Zhang, Chao Chen, Shengshan Hu, Asif Gill, Shirui Pan, 2024
https://scholar.google.com/scholar?q=Large+Language+Model+Watermark+Stealing+With+Mixed+Integer+Programming
25. Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Large Reasoning Models — Shuliang Liu, Xingyu Li, Hongyi Liu, Dong Fang, Yibo Yan, Bingchen Duan, Qi Zheng, Lingfeng Su, Xuming Hu, 2026
https://scholar.google.com/scholar?q=Distilling+the+Thought%2C+Watermarking+the+Answer%3A+A+Principle+Semantic+Guided+Watermark+for+Large+Reasoning+Models
26. AI Post Transformers: Self-Improving Pretraining With Post-Trained Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-02-self-improving-pretraining-with-post-tra-e37460.mp3
27. AI Post Transformers: Distilling Multi-Agent Reasoning into a Single LLM — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-22-distilling-multi-agent-reasoning-into-a-143263.mp3
28. AI Post Transformers: AgenticQwen and Small Industrial Tool Agents — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-27-agenticqwen-and-small-industrial-tool-ag-dc676d.mp3
Interactive Visualization: Trace Rewriting Against Unauthorized LLM Distillation