This episode explores Latent Reasoning with Normalizing Flows, a paper that asks whether a standard left-to-right transformer can do its intermediate reasoning in continuous latent states instead of spelling every step out as text. It explains how the method uses a frozen VAE during training to compress written rationales into short latent sequences, then uses shallow normalizing flows so the same autoregressive backbone can predict both latent thought slots and normal answer tokens while preserving exact likelihoods, sampling, and KV-cache-friendly decoding. The discussion highlights matched coding results on Qwen3-8B-Base, where the reported benchmark average rises from 55.8 for the base model to 68.8 for NF-CoT Unified and 70.1 after latent-space reinforcement learning, with strong pass@k gains that suggest better exploration of multiple solution paths. Listeners would find it interesting because it frames latent reasoning as a practical alternative to verbose chain-of-thought, while also noting the current evidence is still narrow, centered on one post-trained coding model and not uniformly better than diffusion baselines on every benchmark.
Sources:
1. Latent Reasoning with Normalizing Flows — Guancheng Tu, Xiangjun Fu, Suhao Yu, Yao Tang, Haoqiang Kang, Lianhui Qin, Yizhe Zhang, Jiatao Gu, 2026
http://arxiv.org/abs/2606.06447
2. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models — Jason Wei, Xuezhi Wang, Denny Zhou, Quoc Le, et al., 2022
https://scholar.google.com/scholar?q=Chain-of-Thought+Prompting+Elicits+Reasoning+in+Large+Language+Models
3. Training Large Language Models to Reason in a Continuous Latent Space — Shibo Hao, Sainbayar Sukhbaatar, Zhiting Hu, Jason Weston, Yuandong Tian, et al., 2024 preprint; COLM 2025
https://scholar.google.com/scholar?q=Training+Large+Language+Models+to+Reason+in+a+Continuous+Latent+Space
4. Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning — Xinghao Chen, Anhao Zhao, Xiaoyu Shen, et al., 2025
https://scholar.google.com/scholar?q=Reasoning+Beyond+Language%3A+A+Comprehensive+Survey+on+Latent+Chain-of-Thought+Reasoning
5. LaDiR: Latent Diffusion Enhances LLMs for Text Reasoning — Haoqiang Kang, Yizhe Zhang, Navdeep Jaitly, Yi-An Ma, Lianhui Qin, et al., 2025
https://scholar.google.com/scholar?q=LaDiR%3A+Latent+Diffusion+Enhances+LLMs+for+Text+Reasoning
6. CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation — Zhenyi Shen, Hanqi Yan, Linhai Zhang, Zhanghao Hu, Yali Du, Yulan He, 2025
https://scholar.google.com/scholar?q=CODI%3A+Compressing+Chain-of-Thought+into+Continuous+Space+via+Self-Distillation
7. Normalizing Flows are Capable Generative Models — Shuangfei Zhai, Ruixiang Zhang, Preetum Nakkiran, David Berthelot, Jiatao Gu, Huangjie Zheng, Tianrong Chen, Miguel Angel Bautista, Navdeep Jaitly, Josh Susskind, 2024
https://scholar.google.com/scholar?q=Normalizing+Flows+are+Capable+Generative+Models
8. Reasoning by Superposition: A Theoretical Perspective on Chain of Continuous Thought — Hanlin Zhu, Shibo Hao, Zhiting Hu, Jiantao Jiao, Stuart Russell, Yuandong Tian, 2025
https://scholar.google.com/scholar?q=Reasoning+by+Superposition%3A+A+Theoretical+Perspective+on+Chain+of+Continuous+Thought
9. Dynamics Within Latent Chain-of-Thought: An Empirical Study of Causal Structure — Zirui Li, Xuefeng Bai, Kehai Chen, Yizhi Li, Jian Yang, Chenghua Lin, Min Zhang, 2026
https://scholar.google.com/scholar?q=Dynamics+Within+Latent+Chain-of-Thought%3A+An+Empirical+Study+of+Causal+Structure
10. Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning — Jingcheng Deng, Zihao Wei, Liang Pang, Junhong Wu, Shicheng Xu, Zenghao Duan, Huawei Shen, 2026
https://scholar.google.com/scholar?q=Latent-GRPO%3A+Group+Relative+Policy+Optimization+for+Latent+Reasoning
11. Chain-of-Thought Matters: Improving Long-Context Language Models with Reasoning Path Supervision — Dawei Zhu et al., 2025
https://arxiv.org/abs/2502.20790
12. Supervised Chain of Thought — Xiang Zhang and Dujian Ding, 2024
https://arxiv.org/abs/2410.14198
13. Large language models can learn and generalize steganographic chain-of-thought under process supervision — Joey Skaf et al., 2025
https://arxiv.org/abs/2506.01926
14. Hybrid Latent Reasoning via Reinforcement Learning — Zhenrui Yue et al., 2025
https://arxiv.org/abs/2505.18454
15. Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach — Jonas Geiping et al., 2025
https://arxiv.org/abs/2502.05171
16. R-KV: Redundancy-aware KV Cache Compression for Reasoning Models — Zefan Cai et al., 2025
https://arxiv.org/abs/2505.24133
17. Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning — Yu Fu et al., 2024
https://arxiv.org/abs/2410.19258
18. AI Post Transformers: Generative Recursive Reasoning in Latent Space — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-21-generative-recursive-reasoning-in-latent-a9371d.mp3
19. AI Post Transformers: MELT: Decoupling Compute From Memory — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-13-melt-decoupling-compute-from-memory-26430c.mp3
20. AI Post Transformers: Reasoning Theater and Unfaithful Chain-of-Thought — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-reasoning-theater-and-unfaithful-chain-o-a4507e.mp3
21. AI Post Transformers: Gradient Descent at Inference Time for LLM Reasoning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-10-gradient-descent-at-inference-time-for-l-20617d.mp3
22. AI Post Transformers: Explicit Information Transmission for Context Compression — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-explicit-information-transmission-for-co-24e3c2.mp3
23. AI Post Transformers: Speculative Decoding in Real vLLM Serving — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-speculative-decoding-in-real-vllm-servin-6f4e2b.mp3
Interactive Visualization: Latent Reasoning with Normalizing Flows