AI Post Transformers

Compressed Convolutional Attention in Latent Space


Listen Later

This episode explores Compressed Convolutional Attention (CCA), a new approach that pushes transformer attention itself into a shared compressed latent space instead of only compressing the KV-cache. It walks through the evolution from standard multi-head attention to MQA, GQA, and MLA, explaining why earlier efficiency methods mostly targeted decode-time memory while leaving much of the attention compute burden intact. The discussion highlights the paper’s main claim that CCA, especially when combined with grouped-query ideas in CCGQA, can reduce parameters, attention FLOPs, and KV-cache size at the same time while outperforming strong baselines in both dense and mixture-of-experts models. Listeners would find it interesting because it gets into the real systems question behind long-context models: whether a cleaner theoretical efficiency idea can actually translate into better economics and practical performance at scale.
Sources:
1. Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space — Tomas Figliolia, Nicholas Alonso, Rishi Iyer, Quentin Anthony, Beren Millidge, 2025
http://arxiv.org/abs/2510.04476
2. Linformer: Self-Attention with Linear Complexity — Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, Hao Ma, 2020
https://scholar.google.com/scholar?q=Linformer%3A+Self-Attention+with+Linear+Complexity
3. Perceiver: General Perception with Iterative Attention — Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, Joao Carreira, 2021
https://scholar.google.com/scholar?q=Perceiver%3A+General+Perception+with+Iterative+Attention
4. DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model — DeepSeek-AI team, 2024
https://scholar.google.com/scholar?q=DeepSeek-V2%3A+A+Strong%2C+Economical%2C+and+Efficient+Mixture-of-Experts+Language+Model
5. Compressed Convolutional Attention: Efficient Attention in a Compressed Latent Space — Tomas Figliolia, Nicholas Alonso, Rishi Iyer, Quentin Anthony, Beren Millidge, 2025
https://scholar.google.com/scholar?q=Compressed+Convolutional+Attention%3A+Efficient+Attention+in+a+Compressed+Latent+Space
6. Fast Transformer Decoding: One Write-Head is All You Need — Noam Shazeer, 2019
https://scholar.google.com/scholar?q=Fast+Transformer+Decoding%3A+One+Write-Head+is+All+You+Need
7. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, Sumit Sanghai, 2023
https://scholar.google.com/scholar?q=GQA%3A+Training+Generalized+Multi-Query+Transformer+Models+from+Multi-Head+Checkpoints
8. StarCoder 2 and The Stack v2: The Next Generation — Anton Lozhkov, Raymond Li, Loubna Ben Allal and many others, 2024
https://scholar.google.com/scholar?q=StarCoder+2+and+The+Stack+v2%3A+The+Next+Generation
9. RoFormer: Enhanced Transformer with Rotary Position Embedding — Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, Yunfeng Liu, 2021
https://scholar.google.com/scholar?q=RoFormer%3A+Enhanced+Transformer+with+Rotary+Position+Embedding
10. Zyda-2: a 5 Trillion Token High-Quality Dataset — Yury Tokpanov, Paolo Glorioso, Quentin Anthony, Beren Millidge, 2024
https://scholar.google.com/scholar?q=Zyda-2%3A+a+5+Trillion+Token+High-Quality+Dataset
11. RazorAttention: Efficient KV Cache Compression through Retrieval Heads — approx. recent efficient inference / attention-systems authors, 2024/2025
https://scholar.google.com/scholar?q=RazorAttention%3A+Efficient+KV+Cache+Compression+through+Retrieval+Heads
12. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — approx. recent LLM inference authors, 2024/2025
https://scholar.google.com/scholar?q=Not+All+Heads+Matter%3A+A+Head-Level+KV+Cache+Compression+Method+with+Integrated+Retrieval+and+Reasoning
13. State-Space Models Can Learn In-Context by Gradient Descent — approx. recent theory / sequence-model authors, 2024/2025
https://scholar.google.com/scholar?q=State-Space+Models+Can+Learn+In-Context+by+Gradient+Descent
14. Structured State Space Models for In-Context Reinforcement Learning — approx. recent sequence-model authors, 2024/2025
https://scholar.google.com/scholar?q=Structured+State+Space+Models+for+In-Context+Reinforcement+Learning
15. SHRP: Specialized Head Routing and Pruning for Efficient Encoder Compression — approx. recent model-compression authors, 2024/2025
https://scholar.google.com/scholar?q=SHRP%3A+Specialized+Head+Routing+and+Pruning+for+Efficient+Encoder+Compression
16. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
17. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
18. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
19. AI Post Transformers: FlashAttention: IO-Aware Fast and Memory-Efficient Attention — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/flashattention-io-aware-fast-and-memory-efficient-attention/
20. AI Post Transformers: Kimi Linear: Efficient Expressive Attention Architecture — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/kimi-linear-efficient-expressive-attention-architecture/
21. AI Post Transformers: RoPE — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/rope/
22. AI Post Transformers: FengHuang for Rack-Scale LLM Inference Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-fenghuang-for-rack-scale-llm-inference-m-62708e.mp3
23. AI Post Transformers: KV Cache TTL for Multi-Turn Agent Scheduling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-kv-cache-ttl-for-multi-turn-agent-schedu-996bf1.mp3
Interactive Visualization: Compressed Convolutional Attention in Latent Space
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof