AI Post Transformers

Atlas: Test-Time Memory for Long Contexts


Listen Later

This episode explores Atlas, a 2025 paper on test-time memorization that asks whether a model with fixed recurrent memory can learn to update that memory during inference and rival Transformers on long-context recall and reasoning. It explains the core tradeoff between Transformer-style KV caches, which preserve near-exact token access at growing cost, and bounded recurrent memory, which must decide what to keep, compress, or forget. The discussion focuses on why earlier recurrent memory systems fell short, then breaks down Atlas's proposed fixes: evaluating memory updates against a window of recent tokens rather than only the newest token, using richer key representations, and learning stronger retention and optimizer-style write rules. Listeners get a clear view of why this matters for post-Transformer architectures, and why fixed-size memory remains both a promising direction and a stubborn bottleneck.
Sources:
1. ATLAS: Learning to Optimally Memorize the Context at Test Time — Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, Vahab Mirrokni, 2025
http://arxiv.org/abs/2505.23735
2. Linear Transformers Are Secretly Fast Weight Programmers — Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber, 2021
https://scholar.google.com/scholar?q=Linear+Transformers+Are+Secretly+Fast+Weight+Programmers
3. Retentive Network: A Successor to Transformer for Large Language Models — Yutao Sun, Li Dong, Shaohan Huang, Furu Wei, et al., 2023
https://scholar.google.com/scholar?q=Retentive+Network%3A+A+Successor+to+Transformer+for+Large+Language+Models
4. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
5. It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization — Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=It%27s+All+Connected%3A+A+Journey+Through+Test-Time+Memorization%2C+Attentional+Bias%2C+Retention%2C+and+Online+Optimization
6. Learning to (Learn at Test Time): RNNs with Expressive Hidden States — Yu Sun et al., 2024
https://scholar.google.com/scholar?q=Learning+to+%28Learn+at+Test+Time%29%3A+RNNs+with+Expressive+Hidden+States
7. RNNs are not Transformers (Yet): The Key Bottleneck on In-Context Retrieval — Kaiyue Wen, Xingyu Dang, Kaifeng Lyu, 2024
https://scholar.google.com/scholar?q=RNNs+are+not+Transformers+%28Yet%29%3A+The+Key+Bottleneck+on+In-Context+Retrieval
8. Test-time Regression: a Unifying Framework for Designing Sequence Models with Associative Memory — Ke Alexander Wang, Jiaxin Shi, Emily B. Fox, 2025
https://scholar.google.com/scholar?q=Test-time+Regression%3A+a+Unifying+Framework+for+Designing+Sequence+Models+with+Associative+Memory
9. BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack — Yuri Kuratov et al., 2024
https://scholar.google.com/scholar?q=BABILong%3A+Testing+the+Limits+of+LLMs+with+Long+Context+Reasoning-in-a-Haystack
10. SCBench: A KV Cache-Centric Analysis of Long-Context Methods — Yucheng Li et al., 2024
https://scholar.google.com/scholar?q=SCBench%3A+A+KV+Cache-Centric+Analysis+of+Long-Context+Methods
11. Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks — Zheng Wang et al., 2024
https://scholar.google.com/scholar?q=Model+Tells+You+Where+to+Merge%3A+Adaptive+KV+Cache+Merging+for+LLMs+on+Long-Context+Tasks
12. Forgetting Transformer: Softmax Attention with a Forget Gate — Zhixuan Lin et al., 2025
https://scholar.google.com/scholar?q=Forgetting+Transformer%3A+Softmax+Attention+with+a+Forget+Gate
13. Test-Time Training Done Right — Tianyuan Zhang et al., 2025
https://scholar.google.com/scholar?q=Test-Time+Training+Done+Right
14. Associative Recurrent Memory Transformer — Ivan Rodkin et al., 2024
https://scholar.google.com/scholar?q=Associative+Recurrent+Memory+Transformer
15. AI Post Transformers: Titans: Learning to Memorize at Test Time — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-20-titans-learning-to-memorize-at-test-time-054662.mp3
16. AI Post Transformers: Gated Delta Networks for Long-Context Retrieval — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-gated-delta-networks-for-long-context-re-706d85.mp3
17. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
18. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
19. AI Post Transformers: Muon Is Scalable for LLM Training — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-muon-is-scalable-for-llm-training-587ed8.mp3
20. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: Atlas: Test-Time Memory for Long Contexts
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof