Plan, Don’t Pose: Long Composite Motion Generation with Text-Aligned BFM
Abstract
Text-to-motion (T2M) generation has broad applications in character animation, virtual avatars, and human-robot interaction. Existing methods typically generate pose trajectories or motion tokens directly from language, forcing a single model to handle semantic interpretation, long-horizon structure, and low-level physical realization. This coupling makes them costly and often unreliable for long, compositional, or semantically dense prompts. We propose Text2BFM, a framework that performs T2M generation directly in the latent policy space of a frozen, pretrained Behavioral Foundation Model (BFM), trained only on paired text–motion data and without video generation or heavy end-to-end motion generators. A text-aligned variational behavioral bottleneck compresses BFM policy-latent sequences into compact continuous latents that are compatible with language and preserve long-horizon behavioral structure. A lightweight conditional flow model generates in this compact space, and the result is decoded into policy latents that drive the frozen BFM in a single closed-loop rollout. Under identical clause boundaries for all methods, Text2BFM raises compositional order accuracy from 0.42 (best kinematic baseline) to 0.67 and reduces boundary discontinuity by 60%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.