PrefixMotion: Prefix-Consistent Conditional Motion Generation
Abstract
Interactive conditional motion generation requires new observations to extend a motion sequence without revising its history. We introduce PrefixMotion, a framework that adapts pretrained motion generators to prefix-consistent streaming inference. Under compatible conditions and coupled randomness, full-sequence and prefix executions agree on their shared outputs. The framework enforces this agreement across attention, shape aggregation, and motion decoding. A clean-motion parameterization of rectified flow preserves compatibility with pretrained weights and geometric supervision while enabling three-evaluation Euler sampling. Instantiated with GENMO, PrefixMotion is evaluated through video-conditioned reconstruction and text-conditioned synthesis. On 120-frame windows from 37 3DPW tracks, four prefix-length comparisons yield maximum tensor deviations below through the global output, given compatible supplied track and camera prefixes. Cached execution is evaluated on 79 complete sequences and accelerates newest-frame updates by 10.6× over window recomputation on the same GPU. PrefixMotion improves all three 3DPW spatial metrics and all three EMDB-2 trajectory metrics over our evaluation of pretrained GENMO. It also achieves lower error on 10 of 13 metrics than OnlineHMR under shared finalized tracks and DROID-SLAM cameras, with model-specific EMDB intrinsics. On HumanML3D, five-trial mean FID improves from 10.169 to 8.110 and R@3 from 0.475 to 0.492. These results show that explicit temporal constraints can support efficient streaming while preserving conditional generation and improving selected reconstruction metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.