VestEdit: Inversion-Free Video Editing via Spatio-Temporal Transport Stabilization
Abstract
Inversion-free flow editing modifies videos directly through a pretrained backbone's velocity field, avoiding costly noise inversion. However, naive frame-wise velocity construction can entangle the spatial support of an edit with motion across video frames. The resulting residual may leak into non-edited regions or follow directions that are inconsistent with source-frame correspondences, causing flicker, boundary deformation, and background drift. We introduce VestEdit, a training-free correction layer inserted between velocity construction and ODE integration. A semantic-topological spatial projection suppresses residuals outside the editable support, while a norm-preserving spectral projection redirects the spatially corrected field toward an attention-induced low-frequency subspace. On the six FiVE tasks, VestEdit achieves a FiVE-Acc of 50.13 and obtains the best reported six-task average text-alignment scores, with CLIP-S of 27.71 and CLIP-E of 27.87, without additional backbone forward evaluations. These results indicate that explicit spatial and spectral conditioning of the editing field can improve the preservation–alignment trade-off of inversion-free video editing. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.