ReMoVe: Refining a Video Model's Own Motion for Coherent Human Video Generation
Abstract
Video generation models have become increasingly capable of synthesizing realistic human motion, yet they often lose track of the very structure that makes the motion plausible. As generated human motion becomes more complex, videos can suffer from progressive body deformation and appearance drift across frames. Existing motion-controlled methods use explicit representations to stabilize generation, but typically rely on pre-defined motion sequences as references. This raises a fundamental question: if the base generator already produces a plausible coarse motion, why replace it with an external one? **ReMoVe** answers this question by treating the coarse output as a motion hypothesis rather than a final result. ReMoVe recovers and refines the underlying motion, then renders the refined motion as 3D mesh and transfers facial appearance onto it, thereby improving motion accuracy and identity consistency. We further derive depth from the same 3D scene and camera trajectory to impose scene-level spatial consistency beyond the mesh.By jointly conditioning video generation models on mesh and depth, ReMoVe synthesizes high-quality motion videos with controllable camera trajectories and enhanced spatio-temporal consistency. Extensive experiments demonstrate that ReMoVe improves human structural stability, appearance consistency, temporal coherence, and motion-following fidelity, while retaining the base generator's motion prior and offering accurate camera control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.