DyCoPart: Dynamic Body-Part Coordination for Text-to-Motion Generation
Abstract
Text-to-motion (T2M) generation aims to synthesize realistic and physically plausible human motion sequences that are semantically consistent with natural-language descriptions. Despite substantial progress, existing T2M methods still face two key limitations: holistic models tend to miss fine-grained local articulations, while fixed part-based models struggle to capture motion-dependent body coordination due to predefined body-part partitions. To address these limitations, we propose DyCoPart, a dynamic coordinated part-aware framework for text-to-motion generation. DyCoPart first encodes joint-level kinematic dependencies with a lightweight graph module and softly routes joints into six dynamic parts under anatomical-prior guidance, producing part tokens that capture both fine-grained local dynamics and motion-dependent coordination. These part tokens are then processed by a shared spatial Part State Space Model, through which whole-body spatial structure is further modeled in the dynamic-part space. Based on the coordinated part representations, a directed causal coordination rectifier is introduced, where temporally causal cross-part messages are propagated among dynamic-part slots so that whole-body coherence can be recovered from corrupted or misaligned local dynamics.Experiments on HumanML3D and KIT-ML demonstrate strong quantitative performance, and coherence-level evaluations further show that DyCoPart generates more fine-grained, physically plausible, and globally coordinated whole-body motions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.