This episode explores TIDE, a transformer variant that lets every layer re-access the original token identity instead of relying entirely on contextual hidden states to preserve that information. It explains how this design targets rare-token failures and “contextual collapse,” where tokens appearing in similar contexts can become too hard for the model to distinguish, especially in scientific, biomedical, or code-heavy text. The discussion walks through TIDE’s mechanism of token-indexed memory tables and layer-wise routing, framing it as a lightweight side channel rather than retrieval or mixture-of-experts. Listeners would find it interesting because it gets at a basic but rarely questioned assumption in modern transformers and asks whether a small architectural change could improve how models handle the long tail of language.
Sources:
1. TIDE: Every Layer Knows the Token Beneath the Context — Ajay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang, Mehrdad Farajtabar, Minsik Cho, 2026
http://arxiv.org/abs/2605.06216
2. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones and others, 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
3. Transformer Feed-Forward Layers Are Key-Value Memories — Mor Geva, Roei Schuster, Jonathan Berant, Omer Levy, 2021
https://scholar.google.com/scholar?q=Transformer+Feed-Forward+Layers+Are+Key-Value+Memories
4. Locating and Editing Factual Associations in GPT — Kevin Meng, David Bau, Alex Andonian, Yonatan Belinkov, 2022
https://scholar.google.com/scholar?q=Locating+and+Editing+Factual+Associations+in+GPT
5. TIDE: Every Layer Knows the Token Beneath the Context — Ajay Jaiswal, Lauren Hannah, Han-Byul Kim, Duc Hoang, Mehrdad Farajtabar, Minsik Cho, 2026
https://scholar.google.com/scholar?q=TIDE%3A+Every+Layer+Knows+the+Token+Beneath+the+Context
6. Adaptive Input Representations for Neural Language Modeling — Alexei Baevski, Michael Auli, 2018
https://scholar.google.com/scholar?q=Adaptive+Input+Representations+for+Neural+Language+Modeling
7. CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation — Jonathan H. Clark, Dan Garrette, Iulia Turc, John Wieting, 2021
https://scholar.google.com/scholar?q=CANINE%3A+Pre-training+an+Efficient+Tokenization-Free+Encoder+for+Language+Representation
8. ByT5: Towards a token-free future with pre-trained byte-to-byte models — Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou, Sharan Narang, Mihir Kale, Adam Roberts, Colin Raffel, 2021
https://scholar.google.com/scholar?q=ByT5%3A+Towards+a+token-free+future+with+pre-trained+byte-to-byte+models
9. XLM-V: Overcoming the Vocabulary Bottleneck in Multilingual Masked Language Models — Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, Madian Khabsa, 2023
https://scholar.google.com/scholar?q=XLM-V%3A+Overcoming+the+Vocabulary+Bottleneck+in+Multilingual+Masked+Language+Models
10. Knowledge Neurons in Pretrained Transformers — Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, Furu Wei, 2022
https://scholar.google.com/scholar?q=Knowledge+Neurons+in+Pretrained+Transformers
11. Improving Language Models by Retrieving from Trillions of Tokens — Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford and colleagues, 2022
https://scholar.google.com/scholar?q=Improving+Language+Models+by+Retrieving+from+Trillions+of+Tokens
12. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Patrick Lewis, Ethan Perez, Aleksandara Piktus, Fabio Petroni, Vladimir Karpukhin and colleagues, 2020
https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks
13. Reconsidering Degeneration of Token Embeddings with Definitions for Encoder-Based Pre-Trained Language Models — approx. recent NLP authors; exact list unclear from snippet, recent, likely 2020s
https://scholar.google.com/scholar?q=Reconsidering+Degeneration+of+Token+Embeddings+with+Definitions+for+Encoder-Based+Pre-Trained+Language+Models
14. MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers — approx. recent LLM/memory-systems authors; exact list unclear from snippet, recent, likely 2020s
https://scholar.google.com/scholar?q=MemoryLLM%3A+Plug-n-Play+Interpretable+Feed-Forward+Memory+for+Transformers
15. The FFN as a Key-Value Memory: Functional Specialization in Transformer Computation — approx. recent mechanistic-interpretability authors; exact list unclear from snippet, recent, likely 2020s
https://scholar.google.com/scholar?q=The+FFN+as+a+Key-Value+Memory%3A+Functional+Specialization+in+Transformer+Computation
16. MrT5: Dynamic Token Merging for Efficient Byte-Level Language Models — approx. recent efficiency/LM authors; exact list unclear from snippet, recent, likely 2020s
https://scholar.google.com/scholar?q=MrT5%3A+Dynamic+Token+Merging+for+Efficient+Byte-Level+Language+Models
17. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
18. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
19. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
20. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
21. AI Post Transformers: Long Short-Term Memory and Vanishing Gradients — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-long-short-term-memory-and-vanishing-gra-72448c.mp3
Interactive Visualization: TIDE and the Rare Token Problem