AI Post Transformers

Training Million-Token LLMs Beyond the Memory Barrier


Listen Later

This episode explores how the OOMB training system tries to break the memory bottleneck that makes million-token language model training impractical, focusing on why training long contexts is much harder than simply extending inference-time context windows. It explains the paper’s core ideas in plain language, including chunk-recurrent training that recomputes activations during backpropagation, O(1)-style activation memory, and the harder remaining problem of storing and moving the KV cache across extremely long sequences. The discussion also weighs the paper’s central claim with healthy skepticism, asking whether fitting multi-million-token training steps on a single GPU proves genuinely useful long-range learning or mainly demonstrates a strong systems optimization. Listeners would find it interesting because it connects deep learning mechanics, hardware limits, and competing long-context strategies like Ring Attention into a clear debate about what real progress in long-context LLMs should look like.
Sources:
1. Out of the Memory Barrier: A Highly Memory Efficient Training System for LLMs with Million-Token Contexts — Wenhao Li, Daohai Yu, Gen Luo, Yuxin Zhang, Fei Chao, Rongrong Ji, Yifan Wu, Jiaxin Liu, Ziyang Gong, Zimu Liao, 2026
http://arxiv.org/abs/2602.02108
2. Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context — Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V. Le, Ruslan Salakhutdinov, 2019
https://scholar.google.com/scholar?q=Transformer-XL%3A+Attentive+Language+Models+Beyond+a+Fixed-Length+Context
3. Recurrent Memory Transformer — Aydar Bulatov, Yuri Kuratov, Mikhail S. Burtsev, 2022
https://scholar.google.com/scholar?q=Recurrent+Memory+Transformer
4. Ring Attention with Blockwise Transformers for Near-Infinite Context — Hao Liu, Matei Zaharia, Pieter Abbeel, 2023
https://scholar.google.com/scholar?q=Ring+Attention+with+Blockwise+Transformers+for+Near-Infinite+Context
5. Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention — Tsendsuren Munkhdalai, Manaal Faruqui, Siddharth Gopal, 2024
https://scholar.google.com/scholar?q=Leave+No+Context+Behind%3A+Efficient+Infinite+Context+Transformers+with+Infini-attention
6. LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models — Yingfeng Chen et al., 2024
https://scholar.google.com/scholar?q=LongLoRA%3A+Efficient+Fine-tuning+of+Long-Context+Large+Language+Models
7. Efficient Memory Management for Large Language Model Serving with PagedAttention — Woosuk Kwon et al., 2023
https://scholar.google.com/scholar?q=Efficient+Memory+Management+for+Large+Language+Model+Serving+with+PagedAttention
8. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism — Mohammad Shoeybi et al., 2020
https://scholar.google.com/scholar?q=Megatron-LM%3A+Training+Multi-Billion+Parameter+Language+Models+Using+Model+Parallelism
9. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie et al., 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
10. ZeRO-Offload: Democratizing Billion-Scale Model Training — Samyam Rajbhandari et al., 2020
https://scholar.google.com/scholar?q=ZeRO-Offload%3A+Democratizing+Billion-Scale+Model+Training
11. Long Context Compression with Activation Beacon — approx. Liu et al., 2024
https://scholar.google.com/scholar?q=Long+Context+Compression+with+Activation+Beacon
12. Boosting Long-Context Information Seeking via Query-Guided Activation Refilling — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=Boosting+Long-Context+Information+Seeking+via+Query-Guided+Activation+Refilling
13. Kvlink: Accelerating Large Language Models via Efficient KV Cache Reuse — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=Kvlink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
14. SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=SuperOffload%3A+Unleashing+the+Power+of+Large-Scale+LLM+Training+on+Superchips
15. SPPO: Efficient Long-Sequence LLM Training via Adaptive Sequence Pipeline Parallel Offloading — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=SPPO%3A+Efficient+Long-Sequence+LLM+Training+via+Adaptive+Sequence+Pipeline+Parallel+Offloading
16. Keep the Cost Down: A Review on Methods to Optimize LLM's KV-Cache Consumption — approx. unknown from snippet, 2024 or 2025
https://scholar.google.com/scholar?q=Keep+the+Cost+Down%3A+A+Review+on+Methods+to+Optimize+LLM%27s+KV-Cache+Consumption
17. AI Post Transformers: DeepSeek-V4 and Practical Million-Token Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-25-deepseek-v4-and-practical-million-token-6f4de1.mp3
18. AI Post Transformers: KVSwap for Disk-Aware Long-Context On-Device Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-kvswap-for-disk-aware-long-context-on-de-f3c15e.mp3
19. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3
20. AI Post Transformers: CacheFlow and 3D-Parallel KV Cache Restoration — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-01-cacheflow-and-3d-parallel-kv-cache-resto-8db883.mp3
21. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
22. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
Interactive Visualization: Training Million-Token LLMs Beyond the Memory Barrier
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof