This episode explores a mechanistic interpretability paper arguing that transformers often fail at counting not because they lack an internal notion of quantity, but because the pathway that converts that latent count into digit tokens is poorly aligned. It explains key ideas like linear probes, the logit lens, attention, LoRA, constrained next-token evaluation, and autoregressive generation to show how the authors separate “the model knows” from “the model can say.” The discussion highlights striking evidence that intermediate hidden states can encode counts almost perfectly while the corresponding digit readout directions remain nearly orthogonal, creating a readout bottleneck. Listeners would find it interesting because it reframes a familiar model weakness into a precise geometric and causal diagnosis, with implications for how to fix generation failures in modern model families like Pythia, Qwen3, and Mistral.
Sources:
1. Why Transformers Fail at Counting
https://arxiv.org/pdf/2605.03258
2. Teaching Arithmetic to Small Transformers — Andrew McLeish, David Irving, Simon Sokota, Max Black, Berlin Chen, et al., 2024
https://scholar.google.com/scholar?q=Teaching+Arithmetic+to+Small+Transformers
3. Language Models Use Trigonometry to Do Addition — Stephen McLeish, et al., 2024
https://scholar.google.com/scholar?q=Language+Models+Use+Trigonometry+to+Do+Addition
4. Faithfulness of Linear Probes in Transformers — Various probe-critique literature; a representative reference should be cited explicitly by the author, 2019-2024
https://scholar.google.com/scholar?q=Faithfulness+of+Linear+Probes+in+Transformers
5. ROME: Locating and Editing Factual Associations in GPT — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=ROME%3A+Locating+and+Editing+Factual+Associations+in+GPT
6. The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets — Wes Gurnee, et al., 2023
https://scholar.google.com/scholar?q=The+Geometry+of+Truth%3A+Emergent+Linear+Structure+in+Large+Language+Model+Representations+of+True%2FFalse+Datasets
7. A Mathematical Framework for Transformer Circuits — Nelson Elhage, et al., 2021
https://scholar.google.com/scholar?q=A+Mathematical+Framework+for+Transformer+Circuits
8. Finding Transformer Circuits with Edge-Level Attribution Patching — Neel Nanda, et al., 2023
https://scholar.google.com/scholar?q=Finding+Transformer+Circuits+with+Edge-Level+Attribution+Patching
9. Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs — Aaditya K. Singh, DJ Strouse, 2024
https://scholar.google.com/scholar?q=Tokenization+counts%3A+the+impact+of+tokenization+on+arithmetic+in+frontier+LLMs
10. Efficient numeracy in language models through single-token number embeddings — Linus Kreitner, Paul Hager, Jonathan Mengedoht, Georgios Kaissis, Daniel Rueckert, Martin J. Menten, 2025
https://scholar.google.com/scholar?q=Efficient+numeracy+in+language+models+through+single-token+number+embeddings
11. Arithmetic-Based Pretraining Improving Numeracy of Pretrained Language Models — Dominic Petrak, Nafise Sadat Moosavi, Iryna Gurevych, 2023
https://scholar.google.com/scholar?q=Arithmetic-Based+Pretraining+Improving+Numeracy+of+Pretrained+Language+Models
12. Rethinking Weight Tying: Pseudo-Inverse Tying for Stable LM Training and Updates — Jian Gu, Aldeida Aleti, Chunyang Chen, Hongyu Zhang, 2026
https://scholar.google.com/scholar?q=Rethinking+Weight+Tying%3A+Pseudo-Inverse+Tying+for+Stable+LM+Training+and+Updates
13. Latent Causal Probing: A Formal Perspective on Probing with Causal Models of Data — Charles Jin, 2024
https://scholar.google.com/scholar?q=Latent+Causal+Probing%3A+A+Formal+Perspective+on+Probing+with+Causal+Models+of+Data
14. AI Post Transformers: Linear Classifier Probes for Intermediate Layers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-linear-classifier-probes-for-intermediat-927ae3.mp3
15. AI Post Transformers: Neural Chameleons and Evading Activation Monitors — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-neural-chameleons-and-evading-activation-bc470e.mp3
16. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
17. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
18. AI Post Transformers: Latent Space as a New Computational Paradigm — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-05-latent-space-as-a-new-computational-para-810f39.mp3
Interactive Visualization: Why Transformers Fail at Counting