AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities
Abstract
While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate visual transformations on silent clips or isolated audio editing, leaving complex audio-visual editing and cross-modal consistency underexplored. We introduce **AVE-Compass**, a comprehensive benchmark with 145 curated source videos, 196 human-verified editing instructions, and 2,688 fine-grained checklist items. It evaluates Instruction Following, Fidelity Preservation, Realism, and Editing Intent through checklist-based MLLM judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics. Evaluation of six systems shows persistent failures in cross-modal execution and non-target preservation. We further propose **AVE-Agent**, which decomposes instructions into dependent subtasks and improves edits through an evaluator-guided loop. AVE-Agent improves instruction execution, Fidelity Preservation, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.