acceptodds
Under review as a conference paper at ICLR 2027

VMEval: A Multi-Dimensional and Interpretable Framework for Video-Music Harmony Evaluation

Abstract

Recent advances in audio-visual learning have primarily focused on cross-modal generation tasks such as video-to-music generation. However, the field still lacks a systematic and objective framework for evaluating video-music harmony, particularly one that simultaneously supports quantitative scoring and qualitative interpretation. Existing approaches are limited to either global compatibility estimation for cross-modal retrieval or descriptive analysis with multimodal large language models (MLLMs) without targeted supervision for harmony evaluation. In this work, we propose VMEval, a multi-dimensional and interpretable framework for video-music harmony evaluation. We decompose video-music harmony into six structured features including rhythm, emotion, dynamics, structure, atmosphere, and content, which provide complementary views of cross-modal consistency and collectively cover the major perceptual aspects. Based on these features, VMEval learns fine-grained cross-modal consistency and jointly produces quantitative dimension-specific harmony scores and interpretable natural-language explanations. VMEval can be applied to downstream tasks including cross-modal retrieval and the evaluation of video-to-music generation models. Extensive experiments show that VMEval achieves competitive performance on cross-modal retrieval and demonstrates strong agreement with human judgments in evaluating video-music harmony.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.