acceptodds
Under review as a conference paper at ICLR 2027

Image Diffusion Models Are Good 3D Human Motion Generators

Abstract

Text-to-motion generation typically relies on dataset-specific skeletal representations, with models trained separately for benchmarks with different joint layouts. We investigate whether image diffusion architectures can model 3D human motion through a common image-compatible representation that accommodates heterogeneous skeletal layouts while preserving motion-specific kinematic structure. To this end, we organize root-centered joint trajectories into topology-informed joint–time images and model global root dynamics through a complementary latent stream. A motion autoencoder conditions local-motion decoding on the root latent, while a text-conditioned Scalable Interpolant Transformer (SiT) jointly generates both latent streams, with supervision on decoded 3D joint positions further constraining generation in motion space. In addition to dataset-specific models, we train a single motion autoencoder on HumanML3D and KIT-ML, followed by a single text-conditioned SiT shared across both datasets, preserving their respective skeletal definitions without skeleton retargeting to a common joint layout. Experiments on HumanML3D and KIT-ML show that our dataset-specific models achieve state-of-the-art performance, while shared training yields further gains on both benchmarks, demonstrating the benefits of shared modeling across heterogeneous skeletal layouts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.