EditEvolve: Self-Evolving Agents for Context-Aware Video Editing
Abstract
Modern video editing models have acquired increasingly powerful content manipulation capabilities, yet high-quality video editing is not simply an instruction-to-video mapping. The same editing instruction can impose different contextual requirements on different source videos: changes to a subject must coordinate its spatial relations, interactions, and associated effects while preserving unrelated content and temporal continuity. The key challenge therefore lies in understanding the source-video context and organizing an appropriate editing strategy accordingly. Identifying an effective strategy, however, requires understanding not only the scene but also how the available tools and generators behave in practice. To this end, we introduce EditEvolve, a self-evolving framework that learns context-aware video editing strategies from execution experience without updating model parameters. To accommodate diverse contextual requirements, EditEvolve constructs a composable editing workflow space that enables the agent to flexibly organize editing operations. To learn how to effectively use tools and generators, the framework performs offline workflow exploration and feedback-driven revision, turning failures in instruction following, contextual consistency, and content preservation into signals for strategy improvement. EditEvolve further distills execution trajectories into reusable experience that connects contextual conditions with effective editing strategies, allowing prior execution outcomes to guide workflow composition for unseen tasks. Experimental results show that EditEvolve achieves state-of-the-art performance and consistently improves editing performance across different video editing backbones.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.