DeepSeek V4 dropped this week with two open-weights models under MIT license — a 1.6 trillion parameter Pro and a 284 billion parameter Flash, both sporting million-token context windows at a fraction of the compute cost of Western flagships. But the conversation quickly turns to a more subjective question: why does DeepSeek's writing feel warmer, more rhythmic, more vivid than what Claude or GPT produces? This episode unpacks four plausible mechanisms — from a Chinese-heavy pretraining corpus rich in fiction, to domain-expert distillation that preserves stylistic variance, to sampling defaults at temperature 1.0, to an alignment philosophy built on verifiable rewards rather than preference smoothing. We also cover V4's hybrid attention architecture (CSA and HCA), the partial Huawei hardware transition, and the two-stage post-training pipeline that keeps domain experts intact through consolidation. No tidy answers — just the best honest uncertainty we have.
Episode #372461 — open it directly at myweirdprompts.com/372461