AI Post Transformers

Nemotron-TwoTower for Parallel Diffusion Language Modeling


Listen Later

This episode explores Nemotron-TwoTower, an NVIDIA paper that tries to keep the text quality of autoregressive language models while reducing the one-token-at-a-time decoding bottleneck through block-wise diffusion generation. It explains how diffusion language modeling works in practice: future token blocks begin as noisy or masked guesses and are iteratively refined in parallel, rather than emitted strictly one token at a time. The discussion focuses on the paper’s core architectural idea of splitting responsibilities between a frozen pretrained causal context tower and a separate trainable denoiser tower, including layer-aligned cross-attention, reused KV caches and Mamba states, and confidence-based early token commitment. Listeners would find it interesting because it gets beyond benchmark hype and examines the real systems tradeoff the paper is making: higher throughput through heavier refinement steps, balanced against serving complexity, multiple denoising passes, and the risk of losing autoregressive-level reliability.
Sources:
1. Nemotron-TwoTower: Diffusion Language Modeling with Pretrained Autoregressive Context — Fitsum Reda, John Kamalu, Roger Waleffe, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro, 2026
http://arxiv.org/abs/2606.26493
2. Structured Denoising Diffusion Models in Discrete State-Spaces (https://arxiv.org/abs/2107.03006) — Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, Rianne van den Berg, 2021
https://scholar.google.com/scholar?q=Structured+Denoising+Diffusion+Models+in+Discrete+State-Spaces+%28https%3A%2F%2Farxiv.org%2Fabs%2F2107.03006%29
3. Simple and Effective Masked Diffusion Language Models (https://arxiv.org/abs/2406.07524) — Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Volodymyr Kuleshov, 2024
https://scholar.google.com/scholar?q=Simple+and+Effective+Masked+Diffusion+Language+Models+%28https%3A%2F%2Farxiv.org%2Fabs%2F2406.07524%29
4. Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models (https://arxiv.org/abs/2503.09573) — Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang, Volodymyr Kuleshov, 2025
https://scholar.google.com/scholar?q=Block+Diffusion%3A+Interpolating+Between+Autoregressive+and+Diffusion+Language+Models+%28https%3A%2F%2Farxiv.org%2Fabs%2F2503.09573%29
5. Encoder-Decoder Diffusion Language Models for Efficient Training and Inference (https://arxiv.org/abs/2510.22852) — Marianne Arriola, Yair Schiff, Hao Phung, Aaron Gokaslan, Volodymyr Kuleshov, 2025
https://scholar.google.com/scholar?q=Encoder-Decoder+Diffusion+Language+Models+for+Efficient+Training+and+Inference+%28https%3A%2F%2Farxiv.org%2Fabs%2F2510.22852%29
6. Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models — Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, Volodymyr Kuleshov, 2025
https://scholar.google.com/scholar?q=Block+Diffusion%3A+Interpolating+Between+Autoregressive+and+Diffusion+Language+Models
7. Encoder-Decoder Diffusion Language Models for Efficient Training and Inference — Marianne Arriola, Yair Schiff, Hao Phung, Aaron Gokaslan, Volodymyr Kuleshov, 2025
https://scholar.google.com/scholar?q=Encoder-Decoder+Diffusion+Language+Models+for+Efficient+Training+and+Inference
8. Simple and Effective Masked Diffusion Language Models — Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T. Chiu, Alexander Rush, Volodymyr Kuleshov, 2024
https://scholar.google.com/scholar?q=Simple+and+Effective+Masked+Diffusion+Language+Models
9. Large Language Diffusion Models — Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li, 2025
https://scholar.google.com/scholar?q=Large+Language+Diffusion+Models
10. Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data — Jingyang Ou, Shen Nie, Kaiwen Xue, Fengqi Zhu, Jiacheng Sun, Zhenguo Li, Chongxuan Li, 2024
https://scholar.google.com/scholar?q=Your+Absorbing+Discrete+Diffusion+Secretly+Models+the+Conditional+Distributions+of+Clean+Data
11. AI Post Transformers: NeurIPS 2025: Large Language Diffusion Models — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/neurips-2025-large-language-diffusion-models/
12. AI Post Transformers: Nemotron 3 Super Hybrid Mamba-Transformer MoE — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-19-nemotron-3-super-hybrid-mamba-transforme-31ac75.mp3
13. AI Post Transformers: Memory-Bound, Not Bandwidth-Limited Batch-1 LLM Decode — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-02-memory-bound-not-bandwidth-limited-batch-114799.mp3
14. AI Post Transformers: Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/draft-verify-lossless-large-language-model-acceleration-via-self-speculative-dec/
15. AI Post Transformers: JETSPEC and Parallel Tree Speculative Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-27-jetspec-and-parallel-tree-speculative-de-3d144c.mp3
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof