This episode explores whether transformers really need separate query, key, and value projections, treating the problem as weight tying inside attention rather than as a brand-new model design. It explains why KV-cache size and memory bandwidth are major bottlenecks for long-context, on-device decoding, then compares increasingly aggressive sharing schemes, especially the difference between tying keys and values versus tying queries and keys. The discussion emphasizes that the broader sweep happens at 300M parameters, while only the shared-K/V variant is carried to 1.2B scale and remains in contention against practical baselines like grouped-query and multi-query attention. Listeners get a concrete deployment tradeoff: shared K/V can reduce KV-cache memory by about 50 percent at roughly a 3.1 percent perplexity cost, making the episode especially interesting for anyone focused on efficient inference.
Sources:
1. Do Transformers Need Three Projections? Systematic Study of QKV Variants — Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis, 2026
http://arxiv.org/abs/2606.04032v2
2. Attention Is All You Need — Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, 2017
https://arxiv.org/abs/1706.03762
3. Fast Transformer Decoding: One Write-Head is All You Need — Noam Shazeer, 2019
https://arxiv.org/abs/1911.02150
4. GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints — Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, Sumit Sanghai, 2023
https://arxiv.org/abs/2305.13245
5. Do Transformers Need Three Projections? Systematic Study of QKV Variants — Ali Kayyam, Anusha Madan Gopal, M. Anthony Lewis, 2026
https://arxiv.org/abs/2606.04032
6. Using the Output Embedding to Improve Language Models — Ofir Press, Lior Wolf, 2017
https://arxiv.org/abs/1608.05859
7. Linformer: Self-Attention with Linear Complexity — Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, Hao Ma, 2020
https://scholar.google.com/scholar?q=Linformer%3A+Self-Attention+with+Linear+Complexity
8. Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention — Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, Francois Fleuret, 2020
https://scholar.google.com/scholar?q=Transformers+are+RNNs%3A+Fast+Autoregressive+Transformers+with+Linear+Attention
9. Mamba: Linear-Time Sequence Modeling with Selective State Spaces — Albert Gu, Tri Dao, 2023
https://scholar.google.com/scholar?q=Mamba%3A+Linear-Time+Sequence+Modeling+with+Selective+State+Spaces
10. AsymKV: Enabling 1-Bit Quantization of KV Cache with Layer-Wise Asymmetric Quantization Configurations — Qian Tao et al., 2024
https://arxiv.org/abs/2410.13212
11. LongHeads: Multi-Head Attention is Secretly a Long Context Processor — Yi Lu et al., 2024
https://arxiv.org/abs/2402.10685
12. MuDAF: Long-Context Multi-Document Attention Focusing through Contrastive Learning on Attention Heads — Weihao Liu et al., 2025
https://arxiv.org/abs/2502.13963
13. Squeezed Attention: Accelerating Long Context Length LLM Inference — Coleman Hooper et al., 2024
https://arxiv.org/abs/2411.09688
14. Beyond Uniform Query Distribution: Key-Driven Grouped Query Attention — Zohaib Khan et al., 2024
https://arxiv.org/abs/2408.08454
15. Weight Decay Induces Low-Rank Attention Layers — Seijin Kobayashi et al., 2024
https://arxiv.org/abs/2410.23819
16. Dissecting Query-Key Interaction in Vision Transformers — Xu Pan et al., 2024
https://arxiv.org/abs/2405.14880
17. AI Post Transformers: Affordable Large-Scale Decoding Through Model-System Co-Design — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-19-affordable-large-scale-decoding-through-e1d7ed.mp3
18. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
19. AI Post Transformers: Prefill-as-a-Service for Cross-Datacenter KV Cache — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-prefill-as-a-service-for-cross-datacente-7560be.mp3
20. AI Post Transformers: TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-turboquant-online-vector-quantiz-1967b7.mp3
21. AI Post Transformers: How Induction Heads Emerge in Transformers — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-03-how-induction-heads-emerge-in-transforme-a7bfcb.mp3
Interactive Visualization: Do Transformers Need Three Projections?