Better Conditioning, Better Motion: Generation-Oriented Text Representation Learning for Motion Generation
Abstract
Text-to-motion generation aims to synthesize realistic human motion sequences from natural language descriptions. Existing methods typically use a text encoder separately pretrained on proxy objectives. However, such encoders struggle to capture subtle dynamic nuances, leading to inadequate text representations for motion generation. We argue that effective text representations for motion generation should be generation-oriented, i.e., optimized to better support the synthesis of high-fidelity, text-aligned motion sequences. Accordingly, we propose LeGO, a co-optimization framework that Learns Generation-Oriented text representations for motion generation and derives LeGO-CLIP, a motion-specialized text encoder. Specifically, LeGO adapts a pretrained CLIP text encoder into LeGO-CLIP by jointly optimizing it with a motion diffusion model, allowing the encoder to receive differentiable feedback from generation quality through two semantic consistency constraints: a motion consistency constraint and a text consistency constraint. Comprehensive experiments demonstrate the effectiveness of our method from three perspectives: 1) Performance: It achieves state-of-the-art results on HumanML3D, KIT-ML and SnapMoGen benchmarks; 2) Efficiency: It significantly accelerates convergence, requiring up to 8x fewer training iterations; and 3) Transferability: LeGO-CLIP serves as a plug-and-play replacement and consistently improves the performance of diverse motion generation models. The source code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.