SeMoSign: Semantic-Motion-Guided Flow Matching for Sign Language Production
Abstract
Sign language production (SLP) has adopted flow-matching models to synthesize sign motions from spoken-language input. However, existing approaches typically rely on learned flow trajectories without explicitly coordinating semantic representation learning with motion-validity guidance during sampling. This may lead to insufficient language–motion correspondence, trajectory drift, and semantically inconsistent or implausible motions. To address these limitations, we propose SeMoSign, a unified semantic-motion-guided flow-matching framework that promotes consistent generation from latent semantic grounding to trajectory evolution. During training, Semantic-Conditioned Latent Modulation (SCLM) injects textual semantic priors into latent motion representations, establishing structured language–motion correspondence without altering the flow-matching backbone. During inference, Per-step Energy Guidance (PEG) evaluates provisional sampling states with a learned region-aware motion-energy landscape and applies differentiable corrections toward plausible motion configurations at selected steps. By coupling semantic latent guidance with motion-aware trajectory refinement, SeMoSign maintains semantic fidelity while improving the coherence and plausibility of generated motions. Experiments on PHOENIX-2014T demonstrate that SeMoSign improves semantic- and motion-related metrics over representative flow-matching-based SLP methods. Qualitative results further show enhanced semantic consistency, temporal coherence, and motion plausibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.