TopX: Skeleton-Aware In-Context Text-to-Motion Diffusion Across Heterogeneous Topologies
Abstract
Text-to-motion generation has mostly been built for one body: the joint count, joint order and pose parameterisation of a single skeleton are baked into the architecture, so every new character means a new model or a retargeting step. Serving hundreds of production rigs with one network is hard for several reasons: their animations share no common representation, the network must be told what each joint of an unfamiliar tree is, a loss summed over bodies of 34 to 102 joints lets the largest dominate, and a rig outside the training library must be served with no clips or a handful. We present TopX, a text-conditioned motion diffusion transformer that serves 311 heterogeneous animal skeletons with one set of weights, built on three design choices that answer these difficulties. TopX-17, a per-joint representation we design, converts any rigged animation from its kinematic tree, rest pose and motion alone and standardises every (rig, joint, channel) cell by its own statistics, so that rigs of any proportion share one network input without a template, retargeting or inverse kinematics. The skeleton enters as a prompt—its rest pose leads the sequence as a clean frame and its tree distances bias the attention—so nothing in the model is specific to a skeleton. A grouped flow-matching objective with calibrated channel weights keeps every channel group's share of the gradient independent of the number of cells it spans. Scored by a skeleton-aware text–motion evaluator, TopX reaches R@1 0.867, where the real clips score 0.904, and an FID of 0.005, and returns per-joint rotations and a world-space root trajectory that the rig plays directly. For rigs outside the library, the same model has both zero-shot and few-shot capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.