DirectBeat: Behavioral Planning and Rhythmic Execution for Co-Speech Motion
Abstract
Speech-to-motion correspondence is highly ambiguous, making generation difficult to control and inspect. We present DirectBeat, which connects behavioral planning and motion execution through explicit directives. Using pretrained knowledge, speech context, recent memory, and retrieved examples, a frozen multimodal language model predicts two-second behavioral plans. A specialized executor realizes these plans as motion. Warp-VQVAE learns audio-conditioned latent space re-timing via temporal warping to synchronize motion with speech. BEAT2-Directive supplies 101,660 aligned speech–directive–motion triplets from BEAT2 to supervise this interface. DirectBeat achieves FGD scores of 3.631 and 2.76 on the single- and all-speaker benchmarks, reducing FGD by 11% and 38% compared with prior methods. On the unseen TED speech benchmark, it receives the highest aggregate human preference across realism, semantic consistency, synchrony, and diversity among the compared methods. We further demonstrate an application that directives support inspection and persona editing in language without retraining the executor, providing a reusable motion interface for LLM agents. Video results are provided in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.