AI Post Transformers

AllMem for Efficient Long-Context Modeling


Listen Later

This episode explores AllMem, a method for turning pretrained Qwen3 models into long-context systems that keep exact attention over a recent token window while storing older context in a learned memory. It explains why standard transformer attention becomes prohibitively expensive on long chats, books, codebases, and agent traces, and places AllMem in the broader landscape of sliding-window, sparse-attention, recurrent, and memory-augmented architectures. The discussion highlights the paper’s core argument: a hybrid design can preserve sharp local reasoning, compress the distant past through online memory updates, and approach the quality of full attention without the same compute and KV-cache costs. A listener would find it interesting because it connects concrete systems constraints on phones and servers to a specific recipe for making long-context language models more practical.
Sources:
1. AllMem: A Memory-centric Recipe for Efficient Long-context Modeling — Ziming Wang, Xiang Wang, Kailong Peng, Lang Qin, Juan Gabriel Kostelec, Christos Sourmpis, Axel Laborieux, Qinghai Guo, 2026
http://arxiv.org/abs/2602.13680
2. Longformer: The Long-Document Transformer — Iz Beltagy, Matthew E. Peters, Arman Cohan, 2020
https://scholar.google.com/scholar?q=Longformer%3A+The+Long-Document+Transformer
3. Big Bird: Transformers for Longer Sequences — Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Amr Ahmed, et al., 2020
https://scholar.google.com/scholar?q=Big+Bird%3A+Transformers+for+Longer+Sequences
4. Mistral 7B — Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Guillaume Lample, et al., 2023
https://scholar.google.com/scholar?q=Mistral+7B
5. Efficient Streaming Language Models with Attention Sinks — Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, Mike Lewis, 2023
https://scholar.google.com/scholar?q=Efficient+Streaming+Language+Models+with+Attention+Sinks
6. Titans: Learning to Memorize at Test Time — Ali Behrouz, Peilin Zhong, Vahab Mirrokni, 2024
https://scholar.google.com/scholar?q=Titans%3A+Learning+to+Memorize+at+Test+Time
7. Artificial Hippocampus Networks for Efficient Long-Context Modeling — Yunhao Fang, Weihao Yu, Shu Zhong, Qinghao Ye, Xuehan Xiong, Lai Wei, 2025
https://scholar.google.com/scholar?q=Artificial+Hippocampus+Networks+for+Efficient+Long-Context+Modeling
8. The Mamba in the Llama: Distilling and Accelerating Hybrid Models — Junxiong Wang, Daniele Paliotta, Avner May, Alexander M. Rush, Tri Dao, 2024
https://scholar.google.com/scholar?q=The+Mamba+in+the+Llama%3A+Distilling+and+Accelerating+Hybrid+Models
9. MesaNet: Sequence Modeling by Locally Optimal Test-Time Training — Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Agüera y Arcas, João Sacramento, 2025
https://scholar.google.com/scholar?q=MesaNet%3A+Sequence+Modeling+by+Locally+Optimal+Test-Time+Training
10. RULER: What's the Real Context Size of Your Long-Context Language Models? — Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, Boris Ginsburg, 2024
https://scholar.google.com/scholar?q=RULER%3A+What%27s+the+Real+Context+Size+of+Your+Long-Context+Language+Models%3F
11. RazorAttention: Efficient KV Cache Compression Through Retrieval Heads — Hanlin Tang et al., 2024
https://scholar.google.com/scholar?q=RazorAttention%3A+Efficient+KV+Cache+Compression+Through+Retrieval+Heads
12. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024
https://scholar.google.com/scholar?q=Not+All+Heads+Matter%3A+A+Head-Level+KV+Cache+Compression+Method+with+Integrated+Retrieval+and+Reasoning
13. Sliding Window Attention Adaptation — Yijiong Yu et al., 2025
https://scholar.google.com/scholar?q=Sliding+Window+Attention+Adaptation
14. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023
https://scholar.google.com/scholar?q=HyperAttention%3A+Long-context+Attention+in+Near-Linear+Time
15. On-the-Fly Adaptive Distillation of Transformer to Dual-State Linear Attention — Yeonju Ro et al., 2025
https://scholar.google.com/scholar?q=On-the-Fly+Adaptive+Distillation+of+Transformer+to+Dual-State+Linear+Attention
16. Test-Time Training on Nearest Neighbors for Large Language Models — Moritz Hardt and Yu Sun, 2023
https://scholar.google.com/scholar?q=Test-Time+Training+on+Nearest+Neighbors+for+Large+Language+Models
17. Test-Time Learning for Large Language Models — Jinwu Hu et al., 2025
https://scholar.google.com/scholar?q=Test-Time+Learning+for+Large+Language+Models
18. Training Large Reasoning Models Efficiently via Progressive Thought Encoding — Zeliang Zhang et al., 2026
https://scholar.google.com/scholar?q=Training+Large+Reasoning+Models+Efficiently+via+Progressive+Thought+Encoding
19. AI Post Transformers: δ-mem and Online Memory for LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-d-mem-and-online-memory-for-llms-6622fa.mp3
20. AI Post Transformers: MELT: Decoupling Compute From Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-melt-decoupling-compute-from-memory-26430c.mp3
21. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/
22. AI Post Transformers: Ministral 3: Cascade Distillation for Long-Context Multimodal Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-cascade-distillation-for-long-context-mu-0ebd1a.mp3
23. AI Post Transformers: Doc-to-LoRA: Internalizing Context as LoRA — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-29-doc-to-lora-internalizing-context-as-lor-8dd5ec.mp3
24. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
25. AI Post Transformers: Do Transformers Need Three Projections? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-11-do-transformers-need-three-projections-c227d6.mp3
Interactive Visualization: AllMem for Efficient Long-Context Modeling
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof