RVEBench: A Comprehensive Benchmark for Reference-Guided Video Editing
Abstract
Reference-guided video editing is evolving toward complex workflows that combine natural-language instructions with multiple visual references. However, existing benchmarks do not systematically construct challenge-specific cases targeting advanced capabilities and therefore fail to expose capability-specific failures. To address this gap, we introduce RVEBench, a comprehensive and diagnostic benchmark comprising 400 human-calibrated cases and 5,672 fine-grained checklist questions, with six challenging scenarios spanning Spatial Perception, Temporal Modeling, Semantic Reasoning, and Symbolic Fidelity. We further introduce RVE-Eval, a five-dimensional framework that combines sample-specific checklists and rubrics to distinguish edit execution, reference grounding, source preservation, perceptual quality, and physical or logical plausibility. Experiments on ten configurations spanning seven representative models reveal several fundamental limitations. Current systems often produce perceptually polished videos that remain physically or logically implausible. Challenging scenarios expose distinct capability-specific failure patterns. As editing requirements accumulate, instruction execution remains stable, but unintended changes and quality degradation become increasingly severe.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.