SignDubber: Personalized and Synchronized Speech Synthesis from Sign Language Videos
Abstract
Sign language translation has largely focused on converting signing videos into text, while the more practically desired task of synthesizing faithful and expressive speech from sign language remains underexplored. To bridge this gap, we introduce a new task: generating personalized speech that matches the signer’s appearance (timbre), signing rhythm, and visible emotion from the sign video and its translated texts. To tackle this, we present SignDubber, a multimodal conditional speech synthesis framework that jointly models fine‑grained video understanding and cross‑modal alignment for expressive speech generation. Specifically, we adopt a flow-matching-based audio generation model conditioned on text embeddings from SpeechT5 and video features extracted by an audio-synchronization-aware video encoder, achieving effective cross-modal alignment. Since existing datasets mostly provide neutral signing with text annotations and lack paired expressive speech, we design a creative data synthesis pipeline that leverages pre-trained audio‑visual generation models to obtain qualified training pairs. Extensive experiments show that our proposed method generalizes robustly to real‑world videos, accurately capturing identity, rhythm, and emotion. Both evaluation metrics and user studies confirm that our approach significantly outperforms multiple representative baselines, demonstrating its practical value.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.