AI Post Transformers

A Practical Review of Mechanistic Interpretability


Listen Later

This episode explores a review of mechanistic interpretability for transformer language models, focusing on how researchers study internal features, circuits, and claims of universality across models. It explains the core toolkit behind the field, including linear probes, hidden-state analysis, intervention methods, vocabulary projection, and sparse autoencoders, while grounding those ideas in transformer anatomy such as attention heads, MLPs, and the residual stream. The discussion highlights a central tension in the literature: finding information encoded in activations is not the same as proving that information causally drives model behavior, and the episode repeatedly questions where interpretability claims may be overstated. Listeners would find it interesting because it offers a concrete map of a fast-growing area of AI research while also giving a careful critique of the field’s assumptions, evidence, and real-world usefulness.
Sources:
1. A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models — Daking Rai, Yilun Zhou, Shi Feng, Abulhair Saparov, Ziyu Yao, 2024
http://arxiv.org/abs/2407.02646
2. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning — Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Chris Olah, and collaborators, 2023
https://scholar.google.com/scholar?q=Towards+Monosemanticity%3A+Decomposing+Language+Models+With+Dictionary+Learning
3. Sparse Autoencoders Find Highly Interpretable Features in Language Models — Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, Lee Sharkey, 2023
https://scholar.google.com/scholar?q=Sparse+Autoencoders+Find+Highly+Interpretable+Features+in+Language+Models
4. Scaling and Evaluating Sparse Autoencoders — Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, Jeffrey Wu, 2024
https://scholar.google.com/scholar?q=Scaling+and+Evaluating+Sparse+Autoencoders
5. Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders — Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, Janos Kramar, Neel Nanda, 2024
https://scholar.google.com/scholar?q=Jumping+Ahead%3A+Improving+Reconstruction+Fidelity+with+JumpReLU+Sparse+Autoencoders
6. A Mathematical Framework for Transformer Circuits — Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, Chris Olah, 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits
7. Towards Best Practices of Activation Patching in Language Models: Metrics and Methods — Fred Zhang, Neel Nanda, 2024
https://scholar.google.com/scholar?q=Towards+Best+Practices+of+Activation+Patching+in+Language+Models%3A+Metrics+and+Methods
8. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet — Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L. Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, Tom Henighan, 2024
https://scholar.google.com/scholar?q=Scaling+Monosemanticity%3A+Extracting+Interpretable+Features+from+Claude+3+Sonnet
9. Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models — Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, Aaron Mueller, 2025
https://scholar.google.com/scholar?q=Sparse+Feature+Circuits%3A+Discovering+and+Editing+Interpretable+Causal+Graphs+in+Language+Models
10. On the Theoretical Understanding of Identifiable Sparse Autoencoders and Beyond — Jingyi Cui, Qi Zhang, Yifei Wang, Yisen Wang, 2025
https://scholar.google.com/scholar?q=On+the+Theoretical+Understanding+of+Identifiable+Sparse+Autoencoders+and+Beyond
11. Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words — Gouki Minegishi, Hiroki Furuta, Yusuke Iwasawa, Yutaka Matsuo, 2025
https://scholar.google.com/scholar?q=Rethinking+Evaluation+of+Sparse+Autoencoders+through+the+Representation+of+Polysemous+Words
12. Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers — Andrew Nam, Henry Conklin, Yukang Yang, Thomas Griffiths, Jonathan Cohen, Sarah-Jane Leslie, 2025
https://scholar.google.com/scholar?q=Causal+Head+Gating%3A+A+Framework+for+Interpreting+Roles+of+Attention+Heads+in+Transformers
13. Quantifying LLM Attention-Head Stability: Implications for Circuit Universality — Karan Bali, Jack Stanley, Praneet Suresh, Danilo Bzdok, 2026
https://scholar.google.com/scholar?q=Quantifying+LLM+Attention-Head+Stability%3A+Implications+for+Circuit+Universality
14. AI Post Transformers: Mechanistic interpretability: Decoding the AI's Inner Logic: Circuits and Sparse Features — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/mechanistic-interpretability-decoding-the-ais-inner-logic-circuits-and-sparse-fe/
15. AI Post Transformers: Advancing Mechanistic Interpretability with Sparse Autoencoders — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/advancing-mechanistic-interpretability-with-sparse-autoencoders/
16. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
17. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3
18. AI Post Transformers: Internal Safety Collapse in Frontier LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-internal-safety-collapse-in-frontier-llm-8be72f.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof