This episode explores MELT, a looped language model architecture that aims to preserve latent reasoning benefits while preventing KV-cache memory from growing with every reasoning pass. It explains how the paper reframes the problem as a systems and architecture challenge, replacing per-loop cached attention state with a single shared, gated cache per layer inspired by recurrent models like LSTMs and Universal Transformers. The discussion weighs whether this is a genuine shift in reasoning architecture or a narrower engineering improvement, ultimately arguing that the paper’s real contribution is efficient cache management rather than a wholly new paradigm. Listeners would find it interesting for its clear breakdown of inference-time compute scaling, latent reasoning, and why memory bottlenecks could shape the future of practical reasoning models.
Sources:
1. MELT: Decoupling Compute From Memory
https://arxiv.org/pdf/2605.07721
2. Scaling Latent Reasoning via Looped Language Models — Rui-Jie Zhu et al., 2025
https://scholar.google.com/scholar?q=Scaling+Latent+Reasoning+via+Looped+Language+Models
3. Universal Transformers — Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, Lukasz Kaiser, 2018
https://scholar.google.com/scholar?q=Universal+Transformers
4. Reducing Transformer Key-Value Cache Size with Cross-Layer Attention — William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, Jonathan Ragan Kelly, 2024
https://scholar.google.com/scholar?q=Reducing+Transformer+Key-Value+Cache+Size+with+Cross-Layer+Attention
5. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
6. Fast Transformer Decoding: One Write-Head is All You Need — Noam Shazeer, 2019
https://scholar.google.com/scholar?q=Fast+Transformer+Decoding%3A+One+Write-Head+is+All+You+Need
7. Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning — Zeyu Xing, Xing Li, Hui-Ling Zhen, Mingxuan Yuan, Sinno Jialin Pan, 2026
https://scholar.google.com/scholar?q=Beyond+Speedup+--+Utilizing+KV+Cache+for+Sampling+and+Reasoning
8. MemShare: Memory Efficient Inference for Large Reasoning Models through KV Cache Reuse — Kaiwen Chen, Xin Tan, Minchen Yu, Hong Xu, 2025
https://scholar.google.com/scholar?q=MemShare%3A+Memory+Efficient+Inference+for+Large+Reasoning+Models+through+KV+Cache+Reuse
9. Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent Memory — Aydar Bulatov, Yurii Kuratov, Yermek Kapushev, Mikhail Burtsev, 2024
https://scholar.google.com/scholar?q=Beyond+Attention%3A+Breaking+the+Limits+of+Transformer+Context+Length+with+Recurrent+Memory
10. Towards Understanding Distilled Reasoning Models: A Representational Approach — David D. Baek, Max Tegmark, 2025
https://scholar.google.com/scholar?q=Towards+Understanding+Distilled+Reasoning+Models%3A+A+Representational+Approach
11. AI Post Transformers: Gated Linear Attention for Efficient Long Sequences — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-gated-linear-attention-for-efficient-lon-c858ab.mp3
12. AI Post Transformers: PackKV Lossy Compression for KV Caches — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-04-packkv-lossy-compression-for-kv-caches-b37bce.mp3
13. AI Post Transformers: RetrievalAttention for Long-Context LLM Inference — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-17-retrievalattention-for-long-context-llm-ddf566.mp3
14. AI Post Transformers: Parallelizing DeltaNet Linear Transformers over Sequence Length — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-18-parallelizing-deltanet-linear-transforme-2d0377.mp3
15. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
16. AI Post Transformers: Jet-Nemotron and PostNAS for Faster Long Context — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-24-jet-nemotron-and-postnas-for-faster-long-436381.mp3
17. AI Post Transformers: In-Place Test-Time Training for Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-in-place-test-time-training-for-transfor-d0b976.mp3
18. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: MELT: Decoupling Compute From Memory