acceptodds
Under review as a conference paper at ICLR 2027

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

Abstract

As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks both a large-scale human-annotated dataset with complete editing examples and a dedicated reward model for automated assessment. Existing resources are limited by small scale, missing edited outputs, or the absence of human quality labels, while current evaluation often relies on expensive manual inspection or generic vision-language model judges that are not specialized for editing quality. We introduce VEFX-Dataset, a human-annotated dataset containing 5,049 video editing examples spanning 9 major editing categories and 32 subcategories, each labeled along three decoupled dimensions: Instruction Following, Rendering Quality, and Edit Exclusivity. Building on this dataset, we train VEFX-Reward, the first reward model designed specifically for video editing quality assessment. VEFX-Reward jointly processes the source video, the editing instruction, and the edited video, and predicts per-dimension quality scores via ordinal regression. We further release VEFX-Bench, a benchmark of 300 curated video-prompt pairs for a holistic evaluation of video editing models. Experiments show that VEFX-Reward consistently outperforms generic VLM judges and prior reward models on both standard video quality metrics and group-wise preference evaluation, making it an effective tool for benchmarking, model selection, and future reward-driven optimization. We will open-source the dataset, benchmark, and reward model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.