FlowResolve: Constraining Motion and Allocating Guidance in Video Editing
Abstract
Text-driven editing modifies image or video content from natural language instructions. Inversion-free methods do it without mapping the image or video back to a noise latent, steering the sampling trajectory directly with an editing signal. On images this has shown impressive efficiency and structure preservation. Extending it to video remains challenging, and typically fails in three respects. The subject drifts away from the source motion, the background is perturbed, and the edit is not strong enough on the subject. We trace the motion drift to the backbone’s prior, which constrains motion but cannot guarantee consistency with the source motion. The other two failures follow from an editing signal that is misallocated. Based on this analysis we propose FlowResolve, a training-free framework with three components. Multi-scale temporal self-similarity guidance (MTS) constrains motion explicitly. Edit-guided localisation (EGL) confines the editing signal to the region the prompt describes, and reduces the drift at the same time. Target attention focusing (TAF) instead controls the editing strength within that region. Together these mechanisms stabilise the editing signal and add an explicit motion constraint, guiding the flow-based evolution toward the intended target distribution. Extensive experiments on FiVE-Bench show that FlowResolve improves background preservation, motion fidelity and instruction following.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.