Extending the Temporal Horizon of Video Editing via Self-edit Adaptation
Abstract
Recent video editing models are trained on datasets of mostly short videos, which limits their temporal horizon, i.e., the temporal range over which their editing capability is reliably maintained. Although a long video can be processed at once, we find that edits succeeding on short inputs break down into visual artifacts or inaccurate modifications. Editing it in shorter segments instead avoids this breakdown but yields a different edit in each segment, which poses a trade-off between editing quality and temporal consistency. To resolve this, we propose a simple yet effective method, coined , which adapts the model to its own edit that stays within the horizon yet spans the full length of the video. Specifically, we compose consecutive frame groups sampled across the video into a short adaptation clip, covering diverse temporal states despite inter-group discontinuities. Since this discontinuity makes the clip harder to edit even within the horizon, we edit it with source-aligned generation, which keeps the editing trajectory aligned with the source clip. We then train a LoRA on the adaptation clip alone, with its edited latent as the flow-matching target. Finally, the adapted model edits the full video with the same source-aligned generation, avoiding a mismatch between the adaptation target and the inference trajectory. Consequently, without any additional training data, HESA extends the temporal horizon of pretrained video editing models, improving both editing quality and temporal consistency on long and dynamic videos.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.