Image Diffusion Models Are Good 3D Human Motion Generators
Abstract
Text-to-motion generation typically relies on dataset-specific skeletal representations, with models trained separately for benchmarks with different joint layouts. We investigate whether image diffusion architectures can model 3D human motion through a common image-compatible representation that accommodates heterogeneous skeletal layouts while preserving motion-specific kinematic structure. To this end, we organize root-centered joint trajectories into topology-informed joint–time images and model global root dynamics through a complementary latent stream. A motion autoencoder conditions local-motion decoding on the root latent, while a text-conditioned Scalable Interpolant Transformer (SiT) jointly generates both latent streams, with supervision on decoded 3D joint positions further constraining generation in motion space. In addition to dataset-specific models, we train a single motion autoencoder on HumanML3D and KIT-ML, followed by a single text-conditioned SiT shared across both datasets, preserving their respective skeletal definitions without skeleton retargeting to a common joint layout. Experiments on HumanML3D and KIT-ML show that our dataset-specific models achieve state-of-the-art performance, while shared training yields further gains on both benchmarks, demonstrating the benefits of shared modeling across heterogeneous skeletal layouts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.