StrokeForm: Semantic Co-Speech Gesture Generation with Kinematically Grounded Motion Tokens
Abstract
Semantic co-speech gestures express linguistic meaning through body and hand movements. Existing speech-conditioned methods generate natural, rhythmically synchronized motion but remain limited in generating contextually appropriate semantic gestures. Learning such gestures from speech–motion training data is challenging due to their sparse, long-tailed distribution and diverse realizations of the same meaning. Large language models can support context-aware gesture planning, but their plans require an explicit motion interface for reliable execution. We present StrokeForm, a framework that connects language-based gesture planning to motion generation through kinematically grounded motion tokens. An adapted language model plans semantic gestures by specifying their arm configurations, palm orientations, and handshapes, together with contextually relevant linguistic triggers. A grounded tokenizer maps these kinematic attributes to addressable stroke tokens. Anchored near their linguistic triggers, these tokens guide a speech-conditioned masked infiller to generate complete gesture events, while a separate phase condition guides temporal realization. Extensive experiments on BEAT2 demonstrate that StrokeForm achieves higher human-rated semantic appropriateness than the evaluated state-of-the-art baselines and enables controllable realization of specified gesture forms.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.