Metric-Level Universal Kinematic Attribute Control for Text-to-Motion Generation
Abstract
Text-to-motion (T2M) generation has advanced rapidly, yet current systems treat text as a soft semantic specification rather than a metric command: a model can respond to “jump high” but not to “jump 0.4 m”, and writing the number into the prompt has virtually no effect. We propose Metric Attribute Guidance (MAG), a training-free, sampling-time method that equips a frozen vanilla T2M diffusion model with a metric-level control interface over eight universal, body-intrinsic kinematic attributes, seven of which are enforced by differentiable guidance, commanded in absolute physical units alongside the caption. At each late reverse-diffusion step, MAG differentiably recovers world-space joints from the model's posterior mean, extracts soft attribute proxies, and takes a few band-limited, clipped gradient steps toward the commanded values, requiring no fine-tuning and no backbone modification. Applied unchanged to independently trained backbones, MAG raises the mean commanded-realized correlation on HumanML3D from 0.02 to 0.83 under identical commands, with no drop in text fidelity. It transfers to KIT-ML under identical hyper-parameters, reaching centimetre-level median error on the vertical extrema on both benchmarks. These results indicate that metric-level controllability is already latent in standard T2M diffusion models, and can be unlocked purely at sampling time, turning T2M from a semantic sketch into an interface that accepts commands in physical units.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.