AI Post Transformers

ELF and Continuous Language Diffusion


Listen Later

This episode explores ELF: Embedded Language Flows, a continuous-time diffusion language model that stays in embedding space until the final decoding step instead of repeatedly snapping back to discrete tokens during generation. It explains how that design lets the model borrow flow-matching and guidance techniques from image diffusion, while arguing that earlier continuous text models may have underperformed because of token-level constraints rather than any fundamental weakness. The discussion highlights reported results on OpenWebText, where a 105M-parameter ELF model achieves better generative perplexity than 170M baselines with far fewer training tokens and fewer sampling steps, while also extending to translation and summarization. It also digs into the main caveat: whether the gains really come from late discretization and continuous-time modeling, or from a bundle of confounded training and inference choices, making the episode interesting both as a technical walkthrough and as a skeptical evaluation of a bold research claim.
Sources:
1. ELF and Continuous Language Diffusion
https://arxiv.org/pdf/2605.10938
2. Structured Denoising Diffusion Models in Discrete State-Spaces — Jacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow, Rianne van den Berg, 2021
https://scholar.google.com/scholar?q=Structured+Denoising+Diffusion+Models+in+Discrete+State-Spaces
3. Diffusion-LM Improves Controllable Text Generation — Xiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang, Tatsunori B. Hashimoto, 2022
https://scholar.google.com/scholar?q=Diffusion-LM+Improves+Controllable+Text+Generation
4. Simple and Effective Masked Diffusion Language Models — Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin T. Chiu, Alexander M. Rush, Volodymyr Kuleshov, 2024
https://scholar.google.com/scholar?q=Simple+and+Effective+Masked+Diffusion+Language+Models
5. Large Language Diffusion Models — Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li, 2025
https://scholar.google.com/scholar?q=Large+Language+Diffusion+Models
6. Flow Matching for Generative Modeling — Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, Matthew Le, 2023
https://scholar.google.com/scholar?q=Flow+Matching+for+Generative+Modeling
7. Improving and Generalizing Flow-Based Generative Models with Minibatch Optimal Transport — Alexander Tong, Kilian Fatras, Nikolay Malkin, Guillaume Huguet, Yanlei Zhang, Jarrid Rector-Brooks, Guy Wolf, Yoshua Bengio, 2024
https://scholar.google.com/scholar?q=Improving+and+Generalizing+Flow-Based+Generative+Models+with+Minibatch+Optimal+Transport
8. Scaling Rectified Flow Transformers for High-Resolution Image Synthesis — Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Muller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, Robin Rombach, 2024
https://scholar.google.com/scholar?q=Scaling+Rectified+Flow+Transformers+for+High-Resolution+Image+Synthesis
9. Discrete Flow Matching — Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, Yaron Lipman, 2024
https://scholar.google.com/scholar?q=Discrete+Flow+Matching
10. Self-conditioned Embedding Diffusion for Text Generation — Robin Strudel, Corentin Tallec, Florent Altche, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Nikolay Savinov, Sander Dieleman, Laurent Sifre, Remi Leblond, 2022
https://scholar.google.com/scholar?q=Self-conditioned+Embedding+Diffusion+for+Text+Generation
11. Difformer: Empowering Diffusion Models on the Embedding Space for Text Generation — Zhujin Gao, Junliang Guo, Xu Tan, Yongxin Zhu, Fang Zhang, Jiang Bian, Linli Xu, 2022
https://scholar.google.com/scholar?q=Difformer%3A+Empowering+Diffusion+Models+on+the+Embedding+Space+for+Text+Generation
12. LangFlow: Continuous Diffusion Rivals Discrete in Language Modeling — Yuxin Chen, Chumeng Liang, Hangke Sui, Ruihan Guo, Chaoran Cheng, Jiaxuan You, Ge Liu, 2026
https://scholar.google.com/scholar?q=LangFlow%3A+Continuous+Diffusion+Rivals+Discrete+in+Language+Modeling
13. Classifier-Free Diffusion Guidance — Jonathan Ho, Tim Salimans, 2021
https://scholar.google.com/scholar?q=Classifier-Free+Diffusion+Guidance
14. GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models — Alex Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam, Pamela Mishkin, Bob McGrew, Ilya Sutskever, Mark Chen, 2021
https://scholar.google.com/scholar?q=GLIDE%3A+Towards+Photorealistic+Image+Generation+and+Editing+with+Text-Guided+Diffusion+Models
15. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding — Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, S. Sara Mahdavi, Rapha Gontijo Lopes, Tim Salimans, Jonathan Ho, David J. Fleet, Mohammad Norouzi, 2022
https://scholar.google.com/scholar?q=Photorealistic+Text-to-Image+Diffusion+Models+with+Deep+Language+Understanding
16. High-Resolution Image Synthesis with Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2021
https://scholar.google.com/scholar?q=High-Resolution+Image+Synthesis+with+Latent+Diffusion+Models
17. MDLM: Masked Diffusion Language Models — likely the MDLM authors cited as [56] in the paper, 2024
https://scholar.google.com/scholar?q=MDLM%3A+Masked+Diffusion+Language+Models
18. Duo — likely the Duo authors cited as [57] in the paper, 2025
https://scholar.google.com/scholar?q=Duo
19. Latent Diffusion Models — Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Bjorn Ommer, 2022
https://scholar.google.com/scholar?q=Latent+Diffusion+Models
20. LangFlow — the LangFlow authors cited as [10] in the paper, 2026
https://scholar.google.com/scholar?q=LangFlow
21. FLM — the FLM authors cited as [30] in the paper, 2026
https://scholar.google.com/scholar?q=FLM
22. Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution — Aaron Lou, Chenlin Meng, Stefano Ermon, 2024
https://scholar.google.com/scholar?q=Discrete+Diffusion+Modeling+by+Estimating+the+Ratios+of+the+Data+Distribution
23. Scaling Behavior of Discrete Diffusion Language Models — Dimitri von Rutte, Janis Fluri, Omead Pooladzandi, Bernhard Scholkopf, Thomas Hofmann, Antonio Orvieto, 2025
https://scholar.google.com/scholar?q=Scaling+Behavior+of+Discrete+Diffusion+Language+Models
24. Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent Reasoner — Cai Zhou, Chenxiao Yang, Yi Hu, Chenyu Wang, Chubin Zhang, Muhan Zhang, Lester Mackey, Tommi Jaakkola, Stephen Bates, Dinghuai Zhang, 2025
https://scholar.google.com/scholar?q=Coevolutionary+Continuous+Discrete+Diffusion%3A+Make+Your+Diffusion+Language+Model+a+Latent+Reasoner
25. Stay on Topic with Classifier-Free Guidance — Guillaume Sanchez, Honglu Fan, Alexander Spangher, Elad Levi, Pawan Sasanka Ammanamanchi, Stella Biderman, 2023
https://scholar.google.com/scholar?q=Stay+on+Topic+with+Classifier-Free+Guidance
26. Adaptive Classifier-Free Guidance via Dynamic Low-Confidence Masking — Pengxiang Li, Shilin Yan, Joey Tsai, Renrui Zhang, Ruichuan An, Ziyu Guo, Xiaowei Gao, 2025
https://scholar.google.com/scholar?q=Adaptive+Classifier-Free+Guidance+via+Dynamic+Low-Confidence+Masking
27. Studying Classifier(-Free) Guidance From a Classifier-Centric Perspective — Xiaoming Zhao, Alexander G. Schwing, 2025
https://scholar.google.com/scholar?q=Studying+Classifier%28-Free%29+Guidance+From+a+Classifier-Centric+Perspective
28. DEPT: Decoupled Embeddings for Pre-training Language Models — Alex Iacob, Lorenzo Sani, Meghdad Kurmanji, William F. Shen, Xinchi Qiu, Dongqi Cai, Yan Gao, Nicholas D. Lane, 2024
https://scholar.google.com/scholar?q=DEPT%3A+Decoupled+Embeddings+for+Pre-training+Language+Models
29. AI Post Transformers: Generative Modeling via Drifting in One Step — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-05-generative-modeling-via-drifting-in-one-671da0.mp3
30. AI Post Transformers: VL-JEPA for Vision-Language Semantic Prediction — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-12-vl-jepa-for-vision-language-semantic-pre-69c9f4.mp3
31. AI Post Transformers: Why Transformers Fail at Counting — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-why-transformers-fail-at-counting-137924.mp3
32. AI Post Transformers: Mamba-3 for Efficient Sequence Modeling — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-16-mamba-3-for-efficient-sequence-modeling-97a22a.mp3
Interactive Visualization: ELF and Continuous Language Diffusion
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof