acceptodds
Under review as a conference paper at ICLR 2027

ConGRets: Low-Resource, Low-Latency Co-Speech Gesture Synthesis via Contrastive Latent Prediction

Abstract

How small can a complete co-speech gesture system be while scoring on par with state-of-the-art generators? This question matters most for LLM-powered chatbots and dialogue systems, where a response exists as text before speech is synthesized: gestures could be generated without waiting for audio, yet most generators rely on audio the agent does not yet have. ConGRets separates motion and speaker representation learning from text-conditioned gesture prediction. We learn the motion representation contrastively rather than through a VAE or VQ-VAE, then freeze the encoder before decoder training. A separate contrastive encoder derives a global speaker descriptor from reference motion, enabling new-speaker conditioning without generator retraining. A fine-tuned text encoder and fusion head predict the next 10-dimensional motion latent from response text, speaker conditioning, and previous motion, trained by contrastive alignment with the motion space. Thus, one predictor supports memory-based synthesis by selecting and blending recorded units, hybrid synthesis by decoding the selected unit's latent, and direct synthesis by decoding the predicted latent. On 138 common takes from five BEAT speakers excluded from training, and without audio, memory-based synthesis achieves 58% lower position-space Fréchet Gesture Distance (FGD) than the lowest-FGD baseline among seven audio-conditioned systems, with pose diversity close to recorded motion. Hybrid synthesis achieves comparable FGD; direct synthesis achieves lower FGD than three baselines. All modes use at most 12 million generation-time parameters, 27× fewer than that baseline, and after enrollment, loading, and warm-up process the transcript of a 69-second utterance in 15–19 ms on GPU or 23–31 ms on CPU, over 30 times faster than the fastest evaluated baseline on the same input. These results demonstrate that compact text-conditioned prediction can deliver strong distributional gesture quality at low generation cost, providing a practical design for interactive conversational agents and motivating deployment on resource-constrained edge devices.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.