AI Post Transformers

δ-mem and Online Memory for LLMs


Listen Later

This episode explores the paper δ-mem, which argues that long context windows are not the same as true memory and proposes a compact online memory module for frozen language models. It explains how the method uses a tiny mutable state matrix, updated with a delta rule, to store residual errors over time and feed that state back into generation as a low-rank attention correction rather than replaying full conversation history. The discussion also examines why benchmarks like LoCoMo and MemoryAgentBench matter more than generic reasoning tests for evaluating memory, because they probe persistence, conflict resolution, and incremental updating across turns. Listeners would find it interesting because the episode connects an unusual architectural idea to concrete empirical gains, including stronger results on memory-heavy tasks despite using an extremely small memory state.
Sources:
1. $δ$-mem: Efficient Online Memory for Large Language Models — Jingdi Lei, Di Zhang, Junxian Li, Weida Wang, Kaixuan Fan, Xiang Liu, Qihan Liu, Xiaoteng Ma, Baian Chen, Soujanya Poria, 2026
http://arxiv.org/abs/2605.12357
2. Adaptive Switching Circuits — Bernard Widrow, Marcian E. Hoff, 1960
https://isl.stanford.edu/~widrow/papers/c1960adaptiveswitching.pdf
3. Online Learning and Online Convex Optimization — Shai Shalev-Shwartz, 2012
https://cir.nii.ac.jp/crid/1363388845866612864
4. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts — Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, Moritz Hardt, 2020
https://proceedings.mlr.press/v119/sun20b.html
5. TiC-CLIP: Continual Training of CLIP Models — Saurabh Garg, Mehrdad Farajtabar, Hadi Pouransari, Sachin Mehta, Raviteja Vemulapalli, Oncel Tuzel, Vaishaal Shankar, Fartash Faghri, 2024
https://machinelearning.apple.com/research/tic-clip-v2
6. Evaluating Very Long-Term Conversational Memory of LLM Agents — Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, Yuwei Fang, 2024
https://scholar.google.com/scholar?q=Evaluating+Very+Long-Term+Conversational+Memory+of+LLM+Agents
7. Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions — Yuanzhe Hu, Yu Wang, Julian McAuley, 2025
https://scholar.google.com/scholar?q=Evaluating+Memory+in+LLM+Agents+via+Incremental+Multi-Turn+Interactions
8. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2025
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
9. Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory — Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, Deshraj Yadav, 2025
https://scholar.google.com/scholar?q=Mem0%3A+Building+Production-Ready+AI+Agents+with+Scalable+Long-Term+Memory
10. E-mem: Multi-agent based Episodic Context Reconstruction for LLM Agent Memory — Kaixiang Wang, Yidan Lin, Jiong Lou, Zhaojiacheng Zhou, Bunyod Suvonov, Jie Li, 2026
https://scholar.google.com/scholar?q=E-mem%3A+Multi-agent+based+Episodic+Context+Reconstruction+for+LLM+Agent+Memory
11. HiMem: Hierarchical Long-Term Memory for LLM Long-Horizon Agents — Ningning Zhang, Xingxing Yang, Zhizhong Tan, Weiping Deng, Wenyong Wang, 2026
https://scholar.google.com/scholar?q=HiMem%3A+Hierarchical+Long-Term+Memory+for+LLM+Long-Horizon+Agents
12. Continuum Memory Architectures for Long-Horizon LLM Agents — Joe Logan, 2026
https://scholar.google.com/scholar?q=Continuum+Memory+Architectures+for+Long-Horizon+LLM+Agents
13. Gated Delta Networks: Improving Mamba2 with Delta Rule — Songlin Yang, Jan Kautz, Ali Hatamizadeh, 2024/2025
https://scholar.google.com/scholar?q=Gated+Delta+Networks%3A+Improving+Mamba2+with+Delta+Rule
14. Parallelizing Linear Transformers with the Delta Rule over Sequence Length — Songlin Yang, Bailin Wang, Yu Zhang, Yikang Shen, Yoon Kim, 2024
https://scholar.google.com/scholar?q=Parallelizing+Linear+Transformers+with+the+Delta+Rule+over+Sequence+Length
15. Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs — Jonas Hübotter, Sascha Bongni, Ido Hakimi, Andreas Krause, 2024
https://scholar.google.com/scholar?q=Efficiently+Learning+at+Test-Time%3A+Active+Fine-Tuning+of+LLMs
16. Test-Time Learning for Large Language Models — Jinwu Hu, Zhitian Zhang, Guohao Chen, Xutao Wen, Chao Shuai, Wei Luo, Bin Xiao, Yuanqing Li, Mingkui Tan, 2025
https://scholar.google.com/scholar?q=Test-Time+Learning+for+Large+Language+Models
17. Efficient Low Rank Attention for Long-Context Inference in Large Language Models — Tenghui Li, Guoxu Zhou, Xuyang Zhao, Yuning Qiu, Qibin Zhao, 2025
https://scholar.google.com/scholar?q=Efficient+Low+Rank+Attention+for+Long-Context+Inference+in+Large+Language+Models
18. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
19. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
20. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3
21. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
22. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
Interactive Visualization: δ-mem and Online Memory for LLMs
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof