Following complex storylines in long-form TV dramas requires accurately identifying who is speaking each line, a task complicated by multiple characters, unclear audio, and short utterances where voice alone is unreliable. This paper introduces DramaSR-532K, a massive benchmark of 532,000 annotated dialogue lines across 900+ characters, and DramaSR-LRM, a reasoning-based model that combines audio, text, and visual cues via tool-use to attribute speech accurately. The approach notably outperforms existing methods on short utterances. Applications include automated subtitle generation, media indexing, content accessibility tools, and improved recommendation or search systems for video streaming platforms.
Authors: Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu, Jiacheng Shao, Pengfei Chen, Jiannan Ge, Kaiwen Duan, Qi Tian
Paper: https://arxiv.org/abs/2607.02504v1