Text2Cut: Controlling Non-Linear Video Editing with Language
Abstract
Non-linear video editing manipulates clips, tracks, transitions, and effects over structured timelines, yet existing editing systems often rely on software-specific representations, making it challenging to translate natural-language instructions into portable and executable edits. OpenTimelineIO (OTIO) provides a software-agnostic abstraction for timeline-level editing, but does not uniformly represent many effect-level transformations required in practical workflows. To this end, we introduce Text2Cut, a language-driven non-linear video editing framework that extends OTIO with a unified space of executable operations spanning both timeline-level manipulations and effect-level transformations. Rather than directly generating a complete timeline, Text2Cut predicts atomic editing operations that are executed and validated over structured timeline states, with the resulting state compiled into a portable OTIO-based representation. Furthermore, we train Text2Cut on a hybrid corpus of synthetic trajectories and real editing projects, and evaluate it on controlled compositional benchmarks and held-out real editing workflows. Experiments show that Text2Cut achieves highest final-state correctness while maintaining strong execution validity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.