AI Post Transformers

Fast Speech Recognition by Transcript Editing


Listen Later

This episode explores a speech recognition paper that replaces slow left-to-right transcript generation with a faster draft-and-edit approach, where a speech model produces an initial hypothesis and a bidirectional LLM corrects it in parallel. It explains the tradeoff between CTC-based systems, which are fast but weaker at using linguistic context, and autoregressive decoders, which are more expressive but too slow for low-latency use cases like captioning and meetings. The discussion highlights the paper’s key ideas, including transcript editing with insertion slots, latent alignment inspired by CTC, and the use of LoRA to adapt pretrained language models efficiently. Listeners would find it interesting because it shows a concrete path to pushing ASR onto a better speed-accuracy frontier, with reported gains such as a 27x speedup over an autoregressive baseline while staying competitive on word error rate.
Sources:
1. NLE: Non-autoregressive LLM-based ASR by Transcript Editing — Avihu Dekel, Samuel Thomas, Takashi Fukada, George Saon, 2026
http://arxiv.org/abs/2603.08397
2. Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks — Alex Graves, Santiago Fernandez, Faustino Gomez, Jurgen Schmidhuber, 2006
https://scholar.google.com/scholar?q=Connectionist+Temporal+Classification%3A+Labelling+Unsegmented+Sequence+Data+with+Recurrent+Neural+Networks
3. Listen, Attend and Spell — William Chan, Navdeep Jaitly, Quoc V. Le, Oriol Vinyals, 2015
https://scholar.google.com/scholar?q=Listen%2C+Attend+and+Spell
4. Exploring Architectures, Data and Units for Streaming End-to-End Speech Recognition with RNN-Transducer — Hasim Sak, Kanishka Rao, Rohit Prabhavalkar, 2017
https://scholar.google.com/scholar?q=Exploring+Architectures%2C+Data+and+Units+for+Streaming+End-to-End+Speech+Recognition+with+RNN-Transducer
5. Robust Speech Recognition via Large-Scale Weak Supervision — Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever, 2022
https://scholar.google.com/scholar?q=Robust+Speech+Recognition+via+Large-Scale+Weak+Supervision
6. Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict — Yosuke Higuchi, Shinji Watanabe, Nanxin Chen, Tetsuji Ogawa, Tetsunori Kobayashi, 2020
https://scholar.google.com/scholar?q=Mask+CTC%3A+Non-Autoregressive+End-to-End+ASR+with+CTC+and+Mask+Predict
7. Align-Refine: Non-autoregressive Speech Recognition via Iterative Realignment — Ethan A. Chi, Julian Salazar, Katrin Kirchhoff, 2021
https://scholar.google.com/scholar?q=Align-Refine%3A+Non-autoregressive+Speech+Recognition+via+Iterative+Realignment
8. A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond — Yisheng Xiao, Lijun Wu, Junliang Guo, Juntao Li, Min Zhang, Tao Qin, Tie-Yan Liu, 2022
https://scholar.google.com/scholar?q=A+Survey+on+Non-Autoregressive+Generation+for+Neural+Machine+Translation+and+Beyond
9. Encode, Tag, Realize: High-Precision Text Editing — Eric Malmi, Sebastian Krause, Sascha Rothe, Daniil Mirylenka, Aliaksei Severyn, 2019
https://scholar.google.com/scholar?q=Encode%2C+Tag%2C+Realize%3A+High-Precision+Text+Editing
10. Levenshtein Transformer — Jiatao Gu, Changhan Wang, Junbo Zhao, 2019
https://scholar.google.com/scholar?q=Levenshtein+Transformer
11. GECToR - Grammatical Error Correction: Tag, Not Rewrite — Kostiantyn Omelianchuk, Vitaliy Atrasevych, Artem Chernodub, Oleksandr Skurzhanskyi, 2020
https://scholar.google.com/scholar?q=GECToR+-+Grammatical+Error+Correction%3A+Tag%2C+Not+Rewrite
12. FELIX: Flexible Text Editing Through Tagging and Insertion — Jonathan Mallinson, Aliaksei Severyn, Eric Malmi, Guillermo Garrido, 2020
https://scholar.google.com/scholar?q=FELIX%3A+Flexible+Text+Editing+Through+Tagging+and+Insertion
13. LoRA: Low-Rank Adaptation of Large Language Models — Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, 2021
https://scholar.google.com/scholar?q=LoRA%3A+Low-Rank+Adaptation+of+Large+Language+Models
14. Token-and-Duration Transducer — Dmitriy Genzel, Yanzhang He, et al., 2023
https://scholar.google.com/scholar?q=Token-and-Duration+Transducer
15. Mask-Predict: Parallel Decoding of Conditional Masked Language Models — Marjan Ghazvininejad, Omer Levy, Yinhan Liu, Luke Zettlemoyer, 2019
https://scholar.google.com/scholar?q=Mask-Predict%3A+Parallel+Decoding+of+Conditional+Masked+Language+Models
16. QuantSpec: Self-Speculative Decoding with Hierarchical Quantized KV Cache — Rishabh Tiwari, Haocheng Xi, Aditya Tomar, et al., 2025
https://scholar.google.com/scholar?q=QuantSpec%3A+Self-Speculative+Decoding+with+Hierarchical+Quantized+KV+Cache
17. SpeCache: Speculative Key-Value Caching for Efficient Generation of LLMs — Shibo Jie, Yehui Tang, Kai Han, Zhi-Hong Deng, Jing Han, 2025
https://scholar.google.com/scholar?q=SpeCache%3A+Speculative+Key-Value+Caching+for+Efficient+Generation+of+LLMs
18. SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition — Yichong Leng, Xu Tan, Wenjie Liu, Kaitao Song, Rui Wang, Xiang-Yang Li, Tao Qin, Edward Lin, Tie-Yan Liu, 2022
https://scholar.google.com/scholar?q=SoftCorrect%3A+Error+Correction+with+Soft+Detection+for+Automatic+Speech+Recognition
19. Semantic Segmentation with Bidirectional Language Models Improves Long-form ASR — W. Ronny Huang, Hao Zhang, Shankar Kumar, Shuo-Yiin Chang, Tara N. Sainath, 2023
https://scholar.google.com/scholar?q=Semantic+Segmentation+with+Bidirectional+Language+Models+Improves+Long-form+ASR
20. Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR — Yizhou Peng, Hexin Liu, Eng Siong Chng, 2025
https://scholar.google.com/scholar?q=Bi-directional+Context-Enhanced+Speech+Large+Language+Models+for+Multilingual+Conversational+ASR
21. AI Post Transformers: Qwen3.5-Omni Thinker-Talker for Omnimodal Streaming — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-21-qwen35-omni-thinker-talker-for-omnimodal-36b26c.mp3
22. AI Post Transformers: Batch-Aware Expert Routing for Faster MoE Decoding — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-04-batch-aware-expert-routing-for-faster-mo-683ab6.mp3
Interactive Visualization: Fast Speech Recognition by Transcript Editing
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof