Editing Human Motion through Structured Motion Descriptions
Abstract
Text-driven human motion editing aims to modify a source motion sequence according to natural language instructions, enabling users to iteratively refine generated motions. However, current diffusion-based methods denoise the entire motion sequence from random noise within an opaque latent space. This makes it difficult to preserve regions that should remain unchanged and limits the interpretability of the editing process. In this paper, we leverage Structured Motion Descriptions (SMD) to represent skeletal motions as readable text and perform edits directly via a Large Language Model (LLM) guided by instructions. By converting motion into SMD, the LLM selectively modifies only the instruction-relevant regions to achieve precise, localized edits; meanwhile, the readability of SMD enables direct inspection of which joints are altered, over which temporal intervals, and by what magnitude, offering interpretability. We implement this framework through a three-stage training pipeline: supervised fine-tuning to teach the LLM instruction-guided SMD editing, a flow-matching model to enable bidirectional conversion between motion and SMD, and a Group Relative Policy Optimization (GRPO) stage to improve the LLM's numerical accuracy for higher motion fidelity. Extensive experiments on MotionFix demonstrate that our approach achieves state-of-the-art performance while providing interpretability. Code is available in the supplementary materials.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.