Codebook-Geometry Correction: Replacing Test-Time Refinement with a Few Analytic Steps in Trajectory-Controlled Text-to-Motion Generation
Abstract
Text-to-motion generators produce fluent, text-aligned motion, but steering a joint along a prescribed path usually costs either a duplicated control backbone or thousands of per-sample optimization steps at inference. We show that a handful of analytic steps can do the same job, provided the correction is measured in the right geometry. Our central contribution is Codebook-Geometry Correction (CodeGeo): the generator’s part codebooks record which changes its frozen decoder was trained to absorb, and so define a metric in which a correction can be measured; CodeGeo moves a generated clip onto the commanded path by the smallest change that metric allows, in at most eight damped steps in place of 1,600 refinement iterations. CodeGeo needs a start that is already close to the targets, which Key/Value Control (KV-Control) supplies: a thin interface (10.5 M trainable parameters) that feeds the trajectory into a frozen masked transformer as extra key/value memory at every attention layer, over Anatomy-Aware Part Tokenization (PartVQ), one codebook per body part unpacked so that every part-frame token is an addressable attention site. CodeGeo has one hyperparameter, the anisotropy of its step metric, and its setting is not free: it follows how far the commanded anchors spread across body parts, with single-part and body-spanning anchor sets preferring opposite settings in every configuration we tested. On pelvis control the method tracks to 0.24 cm at FID 0.067 in 1/21 the per-sample time of the iterative schedule it replaces; on multi-joint control it reaches 0.25 cm at FID 0.072, improving on MaskControl’s published 0.72 cm and 0.083 on both axes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.