AI Post Transformers

Geometric Memory in Deep Sequence Models


Listen Later

This episode explores whether deep sequence models store knowledge as simple associative lookups or as geometric memories that encode broader relational structure. It discusses a recent paper arguing that, after memorizing graph facts in their weights, sequence models can answer multi-hop path queries as if they were making a much shorter move through embedding space, with learned representations resembling graph-embedding methods like node2vec and DeepWalk. The conversation highlights why that matters mechanistically: it suggests some forms of reasoning may be amortized into the model’s parameters during training rather than reconstructed step by step at inference time. Listeners would find it interesting for its sharp debate over what counts as real reasoning versus a clever shortcut, and for its caution about how far results from synthetic graph settings should generalize to large language models in the wild.
Sources:
1. Deep sequence models tend to memorize geometrically; it is unclear why — Shahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv Kumar, 2025
http://arxiv.org/abs/2510.26745
2. DeepWalk: Online Learning of Social Representations — Bryan Perozzi, Rami Al-Rfou, Steven Skiena, 2014
https://scholar.google.com/scholar?q=DeepWalk%3A+Online+Learning+of+Social+Representations
3. node2vec: Scalable Feature Learning for Networks — Aditya Grover, Jure Leskovec, 2016
https://scholar.google.com/scholar?q=node2vec%3A+Scalable+Feature+Learning+for+Networks
4. Birth of a Transformer: A Memory Viewpoint — Alberto Bietti, Vivien Cabannes, Diane Bouchacourt, Herve Jegou, Leon Bottou, 2023
https://scholar.google.com/scholar?q=Birth+of+a+Transformer%3A+A+Memory+Viewpoint
5. Deep sequence models tend to memorize geometrically; it is unclear why — Shahriar Noroozizadeh, Vaishnavh Nagarajan, Elan Rosenfeld, Sanjiv Kumar, 2025
https://scholar.google.com/scholar?q=Deep+sequence+models+tend+to+memorize+geometrically%3B+it+is+unclear+why
6. The Pitfalls of Next-Token Prediction — Gregor Bachmann, Vaishnavh Nagarajan, 2024
https://scholar.google.com/scholar?q=The+Pitfalls+of+Next-Token+Prediction
7. How Transformers Learn to Plan via Multi-Token Prediction — Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang, 2026
https://scholar.google.com/scholar?q=How+Transformers+Learn+to+Plan+via+Multi-Token+Prediction
8. DeepSeek-V3 Technical Report — DeepSeek-AI and collaborators, 2024
https://scholar.google.com/scholar?q=DeepSeek-V3+Technical+Report
9. Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More — Arvid Frydenlund, 2025
https://scholar.google.com/scholar?q=Language+Models%2C+Graph+Searching%2C+and+Supervision+Adulteration%3A+When+More+Supervision+is+Less+and+How+to+Make+More+More
10. Hopping Too Late: Exploring the Limitations of Large Language Models on Multi-Hop Queries — Eden Biran, Daniela Gottesman, Sohee Yang, Mor Geva, and Amir Globerson, 2024
https://scholar.google.com/scholar?q=Hopping+Too+Late%3A+Exploring+the+Limitations+of+Large+Language+Models+on+Multi-Hop+Queries
11. The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" — Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans, 2024
https://scholar.google.com/scholar?q=The+Reversal+Curse%3A+LLMs+trained+on+%22A+is+B%22+fail+to+learn+%22B+is+A%22
12. In-Context Denoising with One-Layer Transformers: Connections between Attention and Associative Memory Retrieval — Matthew Smart, Alberto Bietti, Anirvan M. Sengupta, 2025
https://scholar.google.com/scholar?q=In-Context+Denoising+with+One-Layer+Transformers%3A+Connections+between+Attention+and+Associative+Memory+Retrieval
13. In-Context Learning as Conditioned Associative Memory Retrieval — Weimin Wu, Teng-Yun Hsiao, Jerry Yao-Chieh Hu, Wenxin Zhang, Han Liu, 2025
https://scholar.google.com/scholar?q=In-Context+Learning+as+Conditioned+Associative+Memory+Retrieval
14. Position-Aware Relational Transformer for Knowledge Graph Embedding — Guangyao Li, Zequn Sun, Wei Hu, Gong Cheng, et al., 2023
https://scholar.google.com/scholar?q=Position-Aware+Relational+Transformer+for+Knowledge+Graph+Embedding
15. Relational Transformer: Toward Zero-Shot Foundation Models for Relational Data — Rishabh Ranjan, Valter Hudovernik, Mark Znidar, Charilaos Kanatsoulis, Roshan Upendra, Mahmoud Mohammadi, Joe Meyer, Tom Palczewski, Carlos Guestrin, Jure Leskovec, 2025
https://scholar.google.com/scholar?q=Relational+Transformer%3A+Toward+Zero-Shot+Foundation+Models+for+Relational+Data
16. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
17. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
Interactive Visualization: Geometric Memory in Deep Sequence Models
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof