TRACE: Typed Reference-Aware Conversational Editing for Humanoid Motion Diffusion
Abstract
Language-conditioned humanoid motion generation turns a textual request into a sequence of robot poses. A short request leaves timing, support, and posture partly unspecified, yet a useful motion must realize these choices coherently and retain them when the user asks for a local change. We present TRACE, a framework for motion generation and reference-aware editing on the Unitree G1 humanoid robot. Our central idea is to express motion structure and requested revisions through typed kinematic quantities that can be grounded in a motion reference. During training, forward kinematics extracts support, root, and hand events from paired motions to supervise a language model. At inference, the language model predicts duration and event intervals from text; their numerical encodings are aligned to motion patches and combined with the original caption in a diffusion Transformer. A gravity-referenced representation directly predicts height and tilt, while trajectory and contact objectives constrain accumulated motion errors. For revision, a typed request specifies a body–time scope and a goal relative to the current motion. A separate diffusion editor uses source features and revised events to realize that goal, with reference projection preserving protected coordinates throughout sampling. Experiments show strong text–motion alignment and human ratings. Matched ablations support structured event conditioning and improved kinematic consistency, while shared-scope comparisons demonstrate more reliable completion of successive edits than the adapted learned editors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.