acceptodds
Under review as a conference paper at ICLR 2027

SignDubber: Personalized and Synchronized Speech Synthesis from Sign Language Videos

Abstract

Sign language translation has largely focused on converting signing videos into text, while the more practically desired task of synthesizing faithful and expressive speech from sign language remains underexplored. To bridge this gap, we introduce a new task: generating personalized speech that matches the signer’s appearance (timbre), signing rhythm, and visible emotion from the sign video and its translated texts. To tackle this, we present SignDubber, a multimodal conditional speech synthesis framework that jointly models fine‑grained video understanding and cross‑modal alignment for expressive speech generation. Specifically, we adopt a flow-matching-based audio generation model conditioned on text embeddings from SpeechT5 and video features extracted by an audio-synchronization-aware video encoder, achieving effective cross-modal alignment. Since existing datasets mostly provide neutral signing with text annotations and lack paired expressive speech, we design a creative data synthesis pipeline that leverages pre-trained audio‑visual generation models to obtain qualified training pairs. Extensive experiments show that our proposed method generalizes robustly to real‑world videos, accurately capturing identity, rhythm, and emotion. Both evaluation metrics and user studies confirm that our approach significantly outperforms multiple representative baselines, demonstrating its practical value.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.