This episode explores DeepSeek-V4’s claim that million-token context windows may finally be practical for real-world use, not just benchmark demos. It explains how the model combines hybrid attention, mixture-of-experts routing, and manifold-constrained hyper-connections to reduce the usual memory and compute costs of long-context transformers while trying to preserve reasoning, coding, and agent performance. The discussion places these design choices in context by comparing them with earlier long-context approaches like Transformer-XL, Longformer, and Big Bird, and by separating headline model size from the more meaningful question of active runtime cost. Listeners would find it interesting because the episode treats the paper not as a simple scaling story, but as a broader systems argument about whether extreme context length can become genuinely useful without hidden tradeoffs.
Sources:
1. DeepSeek-V4 and Practical Million-Token Context
https://cas-bridge.xethub.hf.co/xet-bridge-us/69e864fd6b68f7e6cfc63ca3/4def459c20d33bee897605e5149c7e19d52c49ad592e547dc0ee24044bced2ce?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=cas%2F20260425%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20260425T183145Z&X-Amz-Expires=3600&X-Amz-Signature=f1688b314e79308272cdbfb7c43f2928eb41aa605779c0ce3104d3377e763d53&X-Amz-SignedHeaders=host&X-Xet-Cas-Uid=68488f482f229c24e59d66a0&response-content-disposition=inline%3B+filename*%3DUTF-8%27%27DeepSeek_V4.pdf%3B+filename%3D%22DeepSeek_V4.pdf%22%3B&response-content-type=application%2Fpdf&x-amz-checksum-mode=ENABLED&x-id=GetObject&Expires=1777145505&Policy=eyJTdGF0ZW1lbnQiOlt7IkNvbmRpdGlvbiI6eyJEYXRlTGVzc1RoYW4iOnsiQVdTOkVwb2NoVGltZSI6MTc3NzE0NTUwNX19LCJSZXNvdXJjZSI6Imh0dHBzOi8vY2FzLWJyaWRnZS54ZXRodWIuaGYuY28veGV0LWJyaWRnZS11cy82OWU4NjRmZDZiNjhmN2U2Y2ZjNjNjYTMvNGRlZjQ1OWMyMGQzM2JlZTg5NzYwNWU1MTQ5YzdlMTlkNTJjNDlhZDU5MmU1NDdkYzBlZTI0MDQ0YmNlZDJjZSoifV19&Signature=PduE%7ECXWKurfPe2KmS3YcJQpJcnQmkzyHq%7Ej3dWzYkifRocYIo45kbcjzBIjHTG71uetaQ0rFSPxR27syyAX0bjtIZBIS6d7T7A42ay-uJs0uNjN5mBasT2aQftQBeryDU3bXApQWNVmAxl-kzPcx4WfzWUpZAtFJp4LWZcwLUfBR2Qu%7EZuWg5W6gxJhPKjIfx8MZSdJdzsTz7swo8fX22zwuAxsaMPWU9S-F%7EbNzmvPgcM1OSXm-SZfIzNNmPd3Mi0KcTWf57maRGRzRuDBnkQa4GDj7SdqduSO5g2bfAUzTesQPomRfkqOL1rjvwzcRGLuE1mHgK%7EJxN5EecEqIg__&Key-Pair-Id=K2L8F4GPSG1IFC
2. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models — DeepSeek-AI et al., 2025
https://scholar.google.com/scholar?q=DeepSeek-V3.2%3A+Pushing+the+Frontier+of+Open+Large+Language+Models
3. mHC: Manifold-Constrained Hyper-Connections — Zhenda Xie, Yixuan Wei, Huanqi Cao, Chenggang Zhao, Chengqi Deng, Jiashi Li, Damai Dai, Huazuo Gao, Jiang Chang, Liang Zhao, Shangyan Zhou, Zhean Xu, Zhengyan Zhang, Wangding Zeng, Shengding Hu, Yuqing Wang, Jingyang Yuan, Lean Wang, Wenfeng Liang, 2025
https://scholar.google.com/scholar?q=mHC%3A+Manifold-Constrained+Hyper-Connections
4. Hyper-Connections — Defa Zhu, Hongzhi Huang, Zihao Huang, Yutao Zeng, Yunyao Mao, Banggu Wu, Qiyang Min, Xun Zhou, 2025
https://scholar.google.com/scholar?q=Hyper-Connections
5. Muon is Scalable for LLM Training — Jingyuan Liu et al., 2025
https://scholar.google.com/scholar?q=Muon+is+Scalable+for+LLM+Training
6. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks — Yushi Bai, Shangqing Tu, Jiajie Zhang, Hao Peng, Xiaozhi Wang, Xin Lv, Shulin Cao, Jiazheng Xu, Lei Hou, Yuxiao Dong, Jie Tang, Juanzi Li, 2024
https://scholar.google.com/scholar?q=LongBench+v2%3A+Towards+Deeper+Understanding+and+Reasoning+on+Realistic+Long-context+Multitasks
7. MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers — Chaithanya Bandi, Ben Hertzberg, Geobio Boo, Tejas Polakam, Jeff Da, Sami Hassaan, Manasi Sharma, Andrew Park, Ernesto Hernandez, Dan Rambado, Ivan Salazar, Rafael Cruz, Chetan Rane, Ben Levin, Brad Kenstler, Bing Liu, 2026
https://scholar.google.com/scholar?q=MCP-Atlas%3A+A+Large-Scale+Benchmark+for+Tool-Use+Competency+with+Real+MCP+Servers
8. KVLink: Accelerating Large Language Models via Efficient KV Cache Reuse — Jingbo Yang et al., 2025
https://scholar.google.com/scholar?q=KVLink%3A+Accelerating+Large+Language+Models+via+Efficient+KV+Cache+Reuse
9. HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse — Yuwei An et al., 2025
https://scholar.google.com/scholar?q=HyperRAG%3A+Enhancing+Quality-Efficiency+Tradeoffs+in+Retrieval-Augmented+Generation+with+Reranker+KV-Cache+Reuse
10. ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation — Shihao Wang et al., 2026
https://scholar.google.com/scholar?q=ProphetKV%3A+User-Query-Driven+Selective+Recomputation+for+Efficient+KV+Cache+Reuse+in+Retrieval-Augmented+Generation
11. When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training — Haonan Wang et al., 2024
https://scholar.google.com/scholar?q=When+Precision+Meets+Position%3A+BFloat16+Breaks+Down+RoPE+in+Long-Context+Training
12. LongAttn: Selecting Long-context Training Data via Token-level Attention — Longyun Wu et al., 2025
https://scholar.google.com/scholar?q=LongAttn%3A+Selecting+Long-context+Training+Data+via+Token-level+Attention
13. On the Convergence of Gradient Descent on Learning Transformers with Residual Connections — Zhen Qin et al., 2025
https://scholar.google.com/scholar?q=On+the+Convergence+of+Gradient+Descent+on+Learning+Transformers+with+Residual+Connections
14. ResiDual: Transformer with Dual Residual Connections — Shufang Xie et al., 2023
https://scholar.google.com/scholar?q=ResiDual%3A+Transformer+with+Dual+Residual+Connections
15. HyperAttention: Long-context Attention in Near-Linear Time — Insu Han et al., 2023
https://scholar.google.com/scholar?q=HyperAttention%3A+Long-context+Attention+in+Near-Linear+Time
16. AI Post Transformers: DeepSeek-V3: A Technical Report — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/deepseek-v3-a-technical-report/
17. AI Post Transformers: Memory Sparse Attention for 100M-Token Scaling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-07-memory-sparse-attention-for-100m-token-s-377cff.mp3
18. AI Post Transformers: Ring-linear: Efficient Hybrid Architecture for Long-Context Reasoning — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/ring-linear-efficient-hybrid-architecture-for-long-context-reasoning/
19. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
20. AI Post Transformers: GLM-5: Transitioning from Vibe Coding to Agentic Engineering — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/glm-5-transitioning-from-vibe-coding-to-agentic-engineering/
21. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
Interactive Visualization: DeepSeek-V4 and Practical Million-Token Context