This episode explores how transformers split prediction between knowledge stored in their weights and information inferred from the current prompt, using the paper’s synthetic “bigram world” to make those mechanisms visible. It explains the distinction between global statistical knowledge and true in-context knowledge, then walks through induction heads as a concrete circuit for recalling earlier patterns and continuing them later. The discussion highlights the paper’s main finding that models learn easy dataset-wide averages first, while context-sensitive induction behavior emerges later and requires the right architecture, with two-layer transformers succeeding where one-layer models fail. Listeners would find it interesting because it turns a vague claim about in-context learning into a causal, mechanistic story about how temporary memory may actually form during training.
Sources:
1. Birth of a Transformer: A Memory Viewpoint — Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, Leon Bottou, 2023
http://arxiv.org/abs/2306.00802
2. A Mathematical Framework for Transformer Circuits — Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits
3. In-context Learning and Induction Heads — Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2022
https://scholar.google.com/scholar?q=In-context+Learning+and+Induction+Heads
4. Birth of a Transformer: A Memory Viewpoint — Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, Leon Bottou, 2023
https://scholar.google.com/scholar?q=Birth+of+a+Transformer%3A+A+Memory+Viewpoint
5. What Learning Algorithm Is In-Context Learning? Investigations with Linear Models — Ekin Akyurek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, Denny Zhou, 2023
https://scholar.google.com/scholar?q=What+Learning+Algorithm+Is+In-Context+Learning%3F+Investigations+with+Linear+Models
6. Data Distributional Properties Drive Emergent In-Context Learning in Transformers — Stephanie C. Y. Chan, Adam Santoro, Andrew Lampinen, Jane Wang, Aaditya Singh, Pierre Richemond, Jay McClelland, Felix Hill, 2022
https://scholar.google.com/scholar?q=Data+Distributional+Properties+Drive+Emergent+In-Context+Learning+in+Transformers
7. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
8. Dissecting Recall of Factual Associations in Auto-Regressive Language Models — Mor Geva, Jasmijn Bastings, Katja Filippova, Amir Globerson, 2023
https://scholar.google.com/scholar?q=Dissecting+Recall+of+Factual+Associations+in+Auto-Regressive+Language+Models
9. What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation — Aaditya K. Singh, Ted Moskovitz, Felix Hill, Stephanie C. Y. Chan, Andrew M. Saxe, 2024
https://scholar.google.com/scholar?q=What+needs+to+go+right+for+an+induction+head%3F+A+mechanistic+study+of+in-context+learning+circuits+and+their+formation
10. Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks — Tianyu He, Darshil Doshi, Aritra Das, Andrey Gromov, 2024
https://scholar.google.com/scholar?q=Learning+to+grok%3A+Emergence+of+in-context+learning+and+skill+composition+in+modular+arithmetic+tasks
11. Selective Induction Heads: How Transformers Select Causal Structures In Context — Francesco D'Angelo, Francesco Croce, Nicolas Flammarion, 2025
https://scholar.google.com/scholar?q=Selective+Induction+Heads%3A+How+Transformers+Select+Causal+Structures+In+Context
12. Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models — Shuxun Wang, Qingyu Yin, Chak Tou Leong, Qiang Zhang, Linyi Yang, 2025
https://scholar.google.com/scholar?q=Induction+Head+Toxicity+Mechanistically+Explains+Repetition+Curse+in+Large+Language+Models
13. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
14. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3
15. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
16. AI Post Transformers: Gated Delta Networks for Long-Context Retrieval — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-gated-delta-networks-for-long-context-re-706d85.mp3
Interactive Visualization: How Induction Heads Emerge in Transformers