TimelineBench: Evaluating Video-Editing Agents on Editable Project States
Abstract
In professional video editing, an agent's output is not a rendered video but an editable project that human editors continue to work on. Thus, rigorous evaluation of video-editing agents requires both scalable construction of realistic editing tasks and systematic verification of such project states. We introduce TimelineBench, a benchmark of 781 video-editing tasks that connects task construction and outcome evaluation through a shared structured specification. TimelineBench extracts real workflows from official editing tutorials and organizes them into 48 atomic editing operations and 121 reusable task templates. Each atomic operation specifies both the edit to perform and evaluation rules for the requested changes and required preservation, allowing natural-language instructions and their evaluation to come from the same task specification. TimelineBench evaluates agents directly on the resulting project state, verifying request completion and state preservation, without requiring a reference GUI trajectory or a single gold project. We build OpenCut-TL to run the full benchmark and evaluate seven computer-use agents, whose task success rates range from 3.5% to 87.2%, with preservation failures remaining a major source of error. By unifying realistic task construction with systematic outcome evaluation, TimelineBench provides a scalable and reproducible foundation for evaluating video-editing agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.