HuMoGPT: A Unified Foundation Model for Human Motion Understanding and Synthesis
Abstract
Human motion modeling remains fragmented across specialized tasks, including recognition, captioning, and generation. These tasks are typically addressed by separate models trained on heterogeneous datasets with incompatible representations and supervision, limiting the development of general-purpose motion models. We introduce HuMoGPT, a unified multimodal foundation model that brings diverse human motion understanding and synthesis tasks into a single framework. Our approach has three components. First, we develop MotionRxiv, a large-scale data-curation pipeline that harmonizes heterogeneous motion corpora into a common skeletal representation with consistent normalization, leakage-safe splits, and multitask annotations, yielding approximately M motion clips. We then use these clips to generate M skeleton-grounded question-answering pairs and K multi-turn dialogues. Second, rather than imposing a single representation for understanding and synthesis, we introduce two complementary motion interfaces: SeMoT, a semantic motion tokenizer contrastively aligned with a pretrained language model for understanding, and GeMoT, a continuous temporal autoencoder that preserves geometric fidelity for generation. Third, HuMoGPT integrates these interfaces into a pretrained multimodal language model, using semantic motion representations for understanding and flow matching in the continuous motion latent space for synthesis. Finally, modality-, task-, and domain-routed low-rank adapters enable parameter-efficient multitask learning. HuMoGPT supports captioning, question answering, dialogue, and recognition alongside conditioned generation, completion, in-betweening, and editing. Despite its broad task coverage and single 1.5B backbone, HuMoGPT achieves state-of-the-art motion–text retrieval on HumanML3D, the lowest text-to-motion FID among compared baselines, and captioning performance competitive with a specialized model five times its size.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.