OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing
Abstract
Instruction-based video editing (IVE) is an emerging field with rich application scenarios, yet evaluating the editing models remains a significant challenge. Existing evaluation benchmarks for IVE task suffer from two fundamental limitations. First, task coverage is narrow and largely inherited from image editing, focusing on frame-level spatial manipulations while neglecting video-specific dimensions. Second, current evaluation metrics fail to capture instruction fidelity, allowing models to achieve high scores despite incorrect edits due to the strong visual prior of the original video. To this end, we introduce a comprehensive and clearly structured benchmark for IVE task. Our benchmark systematically decomposes editing tasks along multiple video-specific dimensions, including spatial, temporal, audio, and reference-based editing, going beyond conventional frame-level formulations. Moreover, it explicitly distinguishes between explicit and implicit instructions, incorporating reasoning-based scenarios to better reflect real-world editing requirements. This design enables a more complete and fine-grained evaluation of model capabilities across diverse editing dimensions. Furthermore, we propose a systematic evaluation framework that decomposes editing quality into four complementary dimensions: accuracy, preservation, realism, and consistency with both human judges and state-of-the-art vision language model (VLM). To emphasize the central role of instruction fidelity, we further introduce an accuracy-aware penalty mechanism, where the scores of other dimensions are conditioned on accuracy. This design prevents misleading high scores for visually plausible yet incorrect edits. We conducted experiments evaluating current prominent instruction-based video editing models, comprising both open-source and commercial models. Experimental results reveal that current models remain far from satisfactory in instruction-based video editing. OmniEdit-Bench provides a comprehensive and reliable testbed for instruction-based video editing, offering valuable insights into current model limitations and guiding future research in this direction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.