ForeMotion: Learning Causal Co-Speech Gestures through Future-Informed Supervision
Abstract
Generating co-speech gestures as speech unfolds requires predicting motion from incomplete acoustic and linguistic context. Complete training recordings, however, contain later speech that is unavailable to a causal generator at deployment. The challenge is to use this additional context for learning while preserving the generator's causal speech access at inference. We present ForeMotion, a framework for learning causal co-speech gestures through future-informed gesture supervision. ForeMotion represents motion as discrete gesture tokens and trains a shared autoregressive generator using paired-view gesture distillation. The causal gesture view receives a rolling window of past and current speech, whereas the future-conditioned gesture view accesses additional past and future audio and text from the training clip. Both views predict the same gesture targets from the same permitted motion history. The future-conditioned view produces probability distributions over possible gesture tokens. A gesture-token distillation loss trains the causal view to match these distributions, with gradients stopped through the supervisory predictions. At inference, only the causal view is executed, generating motion autoregressively from available speech and previously generated motion without requiring future speech. Experiments on BEATX demonstrate gains in measured gesture quality over the evaluated causal baselines, with benefits varying across future-context scopes, speaker settings, and evaluation metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.