X-Dancers: Group Choreography as Masked Completion of a Formation-Grounded Group Motion Tensor
Abstract
Music-driven 3D group choreography requires jointly modeling individual motion, evolving formations, global trajectories and their synchronization with music. Existing methods model these factors separately and rely on dedicated mechanisms for different choreographic operations, limiting flexibility and generalization. We present **X-Dancers**, a unified framework that formulates group choreography as *bidirectional masked generation* over a continuous group motion tensor, where each cell couples a dancer's motion representation with its root state to jointly represent motion, formation and trajectory. To learn their joint distribution, we introduce the *ChoreoGraph Transformer*, which uses interleaved temporal and formation-grounded spatial attention to reason over observed group context, while masked conditional diffusion generates the unknown tensor cells. A spatial-temporal-channel masking strategy exposes the model to diverse conditioning patterns, enabling choreography synthesis, long-duration continuation, cast extension and trajectory control via different conditioning masks without task-specific retraining. The experimental results demonstrate its state-of-the-art motion fidelity, music synchronization and formation quality, together with strong zero-shot generalization to long-duration choreography.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.