AI Post Transformers

Hopfield Networks and Transformer Attention as Memory


Listen Later

This episode explores how the 2021 paper “Hopfield Networks Is All You Need” reframes transformer attention as a modern continuous Hopfield network, connecting attention, associative memory, and similarity-based retrieval under one mathematical lens. It explains the core idea of content-addressable memory, contrasts classical binary Hopfield networks with newer differentiable versions, and shows why attention can be understood not just as weighted averaging but as an energy-based retrieval process with fixed points and attractor states. The discussion highlights the paper’s major claims: one-step retrieval, exponential storage capacity, low retrieval error under assumptions, and distinct retrieval regimes such as global averaging and subset averaging. Listeners interested in AI theory will find it compelling because it offers a concrete, less mystical interpretation of transformer heads and suggests practical memory-layer designs grounded in formal guarantees.
Sources:
1. Hopfield Networks is All You Need — Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, Victor Greiff, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter, Sepp Hochreiter, 2020
http://arxiv.org/abs/2008.02217
2. https://isl.stanford.edu/~cover/papers/transIT/0021cove.pdf
https://isl.stanford.edu/~cover/papers/transIT/0021cove.pdf
3. Neural Networks and Physical Systems with Emergent Collective Computational Abilities — John J. Hopfield, 1982
https://scholar.google.com/scholar?q=Neural+Networks+and+Physical+Systems+with+Emergent+Collective+Computational+Abilities
4. A Neural Network with Locality-Sensitive Hashed Dynamics — Dmitry Krotov, John J. Hopfield, 2016
https://scholar.google.com/scholar?q=A+Neural+Network+with+Locality-Sensitive+Hashed+Dynamics
5. Hopfield Networks is All You Need — Hubert Ramsauer, Bernhard Schäfl, Johannes Lehner, Philipp Seidl, Michael Widrich, Thomas Adler, Lukas Gruber, Markus Holzleitner, Milena Pavlović, Geir Kjetil Sandve, Victor Greiff, David Kreil, Michael Kopp, Günter Klambauer, Johannes Brandstetter, Sepp Hochreiter, 2020
https://scholar.google.com/scholar?q=Hopfield+Networks+is+All+You+Need
6. Dense Associative Memory for Pattern Recognition — Dmitry Krotov, John J. Hopfield, 2020
https://scholar.google.com/scholar?q=Dense+Associative+Memory+for+Pattern+Recognition
7. Neural Turing Machines — Alex Graves, Greg Wayne, Ivo Danihelka, 2014
https://scholar.google.com/scholar?q=Neural+Turing+Machines
8. Memory Networks — Jason Weston, Sumit Chopra, Antoine Bordes, 2014
https://scholar.google.com/scholar?q=Memory+Networks
9. End-To-End Memory Networks — Sainbayar Sukhbaatar, Arthur Szlam, Jason Weston, Rob Fergus, 2015
https://scholar.google.com/scholar?q=End-To-End+Memory+Networks
10. Hybrid computing using a neural network with dynamic external memory — Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, Adrià Puigdomènech Badia, Karl Moritz Hermann, Yori Zwols, Georg Ostrovski, Adam Cain, Helen King, Christopher Summerfield, Phil Blunsom, Koray Kavukcuoglu, Demis Hassabis, 2016
https://scholar.google.com/scholar?q=Hybrid+computing+using+a+neural+network+with+dynamic+external+memory
11. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin, 2017
https://scholar.google.com/scholar?q=Attention+Is+All+You+Need
12. Dense Associative Memory is Robust to Adversarial Inputs — Dmitry Krotov, John J. Hopfield, 2018
https://scholar.google.com/scholar?q=Dense+Associative+Memory+is+Robust+to+Adversarial+Inputs
13. A Robust Exponential Associative Memory with Fixed Point Analysis — Mert Demircigil, Judith Heusel, Matthias Löwe, Sven Upgang, Franck Vermet, 2017
https://scholar.google.com/scholar?q=A+Robust+Exponential+Associative+Memory+with+Fixed+Point+Analysis
14. Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks — Juho Lee, Yoonho Lee, Jungtaek Kim, Adam Kosiorek, Seungjin Choi, Yee Whye Teh, 2019
https://scholar.google.com/scholar?q=Set+Transformer%3A+A+Framework+for+Attention-based+Permutation-Invariant+Neural+Networks
15. Deep Sets — Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Ruslan Salakhutdinov, Alexander Smola, 2017
https://scholar.google.com/scholar?q=Deep+Sets
16. Modern Hopfield Networks and Attention for Immune Repertoire Classification — Michael Widrich, Bernhard Schäfl, Milena Pavlović, Hubert Ramsauer, Lukas Gruber, Markus Holzleitner, Geir Kjetil Sandve, Victor Greiff, Sepp Hochreiter, et al., 2020
https://scholar.google.com/scholar?q=Modern+Hopfield+Networks+and+Attention+for+Immune+Repertoire+Classification
17. Attention Heads of Large Language Models — authors unclear from snippet, likely 2024-2025
https://scholar.google.com/scholar?q=Attention+Heads+of+Large+Language+Models
18. Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers — authors unclear from snippet, likely 2024-2025
https://scholar.google.com/scholar?q=Causal+Head+Gating%3A+A+Framework+for+Interpreting+Roles+of+Attention+Heads+in+Transformers
19. Mechanistic Interpretability of Fine-Tuned Vision Transformers on Distorted Images: Decoding Attention Head Behavior for Transparent and Trustworthy AI — authors unclear from snippet, likely 2024-2025
https://scholar.google.com/scholar?q=Mechanistic+Interpretability+of+Fine-Tuned+Vision+Transformers+on+Distorted+Images%3A+Decoding+Attention+Head+Behavior+for+Transparent+and+Trustworthy+AI
20. Iterative Sparse Attention for Long-Sequence Recommendation — authors unclear from snippet, likely recent
https://scholar.google.com/scholar?q=Iterative+Sparse+Attention+for+Long-Sequence+Recommendation
21. An Evolved Universal Transformer Memory — authors unclear from snippet, likely recent
https://scholar.google.com/scholar?q=An+Evolved+Universal+Transformer+Memory
22. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
23. AI Post Transformers: RoPE — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/rope/
24. AI Post Transformers: Jet-Nemotron and PostNAS for Faster Long Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-24-jet-nemotron-and-postnas-for-faster-long-436381.mp3
25. AI Post Transformers: Experimental Comparison of Agentic and Enhanced RAG — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-14-experimental-comparison-of-agentic-and-e-37d8bc.mp3
26. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof