Low-Redundancy Modeling for Latent-Free Human Motion Generation
Abstract
Many recent approaches to human motion generation operate in latent spaces, which require two-stage training and inherit the reconstruction errors of their autoencoders. We revisit generation in raw motion space and identify two key bottlenecks: feature redundancy, which arises from over-parameterized representations, and state redundancy, which arises from invalid off-manifold states and unnatural motions. Building on this observation, we instantiate DirectMotion, a minimalist baseline that combines three simple designs: a compact representation consisting of root features and joint rotations, probability paths constrained to the state space of human motion, and clean-sequence prediction. With only a lightweight DiT network, DirectMotion performs on par with larger latent-space models on the HumanML3D and SnapMoGen datasets. Ablation studies confirm the contribution of removing each form of redundancy. We hope these results encourage the community to reconsider latent-free generation as an inherently simpler paradigm.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.