This episode explores a 2025 paper arguing that decoder-only language models are generically injective on discrete token sequences, meaning their hidden representations can in principle preserve enough information to recover the exact original prompt. It walks through what injectivity and invertibility mean in this setting, why that challenges the common intuition that transformer representations behave like lossy semantic summaries, and how the paper distinguishes this claim from stronger notions of full bijectivity over continuous spaces. The discussion also connects the result to related ideas from normalizing flows, reversible networks, and mechanistic interpretability, while introducing the paper’s constructive recovery method, SipIt. Listeners would find it interesting because the result has unusually sharp implications for both interpretability and privacy: hidden states may be far less abstracted from raw input text than many researchers assume.
Sources:
1. Language Models are Injective and Hence Invertible — Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodolà, 2025
http://arxiv.org/abs/2510.15511
2. Normalizing Flows for Probabilistic Modeling and Inference — George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, Balaji Lakshminarayanan, 2021
https://scholar.google.com/scholar?q=Normalizing+Flows+for+Probabilistic+Modeling+and+Inference
3. The Reversible Residual Network: Backpropagation Without Storing Activations — Aidan N. Gomez, Mengye Ren, Raquel Urtasun, Roger B. Grosse, 2017
https://scholar.google.com/scholar?q=The+Reversible+Residual+Network%3A+Backpropagation+Without+Storing+Activations
4. Sentence Embedding Leaks More Information than You Expect: Generative Embedding Inversion Attack to Recover the Whole Sentence — Haoran Li, Mingshi Xu, Yangqiu Song, 2023
https://scholar.google.com/scholar?q=Sentence+Embedding+Leaks+More+Information+than+You+Expect%3A+Generative+Embedding+Inversion+Attack+to+Recover+the+Whole+Sentence
5. Language Models are Injective and Hence Invertible — Giorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli, Yannis Panagakis, Emanuele Rodola, 2025
https://scholar.google.com/scholar?q=Language+Models+are+Injective+and+Hence+Invertible
6. The non-linear representation dilemma: Is causal abstraction enough for mechanistic interpretability? — Denis Sutter, Julian Minder, Thomas Hofmann, Tiago Pimentel, 2025
https://scholar.google.com/scholar?q=The+non-linear+representation+dilemma%3A+Is+causal+abstraction+enough+for+mechanistic+interpretability%3F
7. On surjectivity of neural networks: Can you elicit any behavior from your model? — Haozhe Jiang, Nika Haghtalab, 2025
https://scholar.google.com/scholar?q=On+surjectivity+of+neural+networks%3A+Can+you+elicit+any+behavior+from+your+model%3F
8. Attention is not all you need: Pure attention loses rank doubly exponentially with depth — Yihe Dong, Jean-Baptiste Cordonnier, Andreas Loukas, 2021
https://scholar.google.com/scholar?q=Attention+is+not+all+you+need%3A+Pure+attention+loses+rank+doubly+exponentially+with+depth
9. Text embeddings reveal (almost) as much as text — John X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. Rush, 2023
https://scholar.google.com/scholar?q=Text+embeddings+reveal+%28almost%29+as+much+as+text
10. Language model inversion — John X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov, Alexander M. Rush, 2023
https://scholar.google.com/scholar?q=Language+model+inversion
11. Better language model inversion by compactly representing next-token distributions — Murtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren, Swabha Swayamdipta, 2025
https://scholar.google.com/scholar?q=Better+language+model+inversion+by+compactly+representing+next-token+distributions
12. Stabilizing Transformer Training by Preventing Attention Entropy Collapse — Shuangfei Zhai et al., 2023
https://scholar.google.com/scholar?q=Stabilizing+Transformer+Training+by+Preventing+Attention+Entropy+Collapse
13. From Condensation to Rank Collapse: A Two-Stage Analysis of Transformer Training Dynamics — Zheng-An Chen and Tao Luo, 2025
https://scholar.google.com/scholar?q=From+Condensation+to+Rank+Collapse%3A+A+Two-Stage+Analysis+of+Transformer+Training+Dynamics
14. Understanding and Minimising Outlier Features in Transformer Training — Bobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag, Thomas Hofmann, 2024
https://scholar.google.com/scholar?q=Understanding+and+Minimising+Outlier+Features+in+Transformer+Training
15. Measuring In-Context Computation Complexity via Hidden State Prediction — Vincent Herrmann, Robert Csordas, Jurgen Schmidhuber, 2025
https://scholar.google.com/scholar?q=Measuring+In-Context+Computation+Complexity+via+Hidden+State+Prediction
16. Transformers without Normalization — Jiachen Zhu, Xinlei Chen, Kaiming He, Yann LeCun, Zhuang Liu, 2025
https://scholar.google.com/scholar?q=Transformers+without+Normalization
17. AI Post Transformers: Language Models are Injective and Hence Invertible — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-21-language-models-are-injective-an-7545e0.mp3
18. AI Post Transformers: RoPE — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/rope/
19. AI Post Transformers: Mistral 7B: Superior Performance in a Smaller Package — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/mistral-7b-superior-performance-in-a-smaller-package/
20. AI Post Transformers: ALiBi: Attention with Linear Biases Enables Length Extrapolation — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/alibi-attention-with-linear-biases-enables-length-extrapolation/
21. AI Post Transformers: CacheSlide: Position-Aware KV Cache Reuse for Agent LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-cacheslide-position-aware-kv-cache-reuse-cd59c7.mp3
Interactive Visualization: Episode: Language Models are Injective and Hence Invertible