AUDIO-COMPASS: REFERENCE-FREE COMPONENT-WISE EVALUATOR FOR AUDIO–TEXT SEMANTIC ALIGNMENT
Abstract
Many audio–text evaluation metrics summarize semantic alignment with a single overall score, motivating explicit assessment along consistent semantic dimensions. We introduce \model, a 30B reference-free audio judge that produces overall and component-wise scores for sound events, sources, acoustic attributes, and temporal relations, along with structured explanations of matched, missing, and incorrect information. To train \model, we construct \dataset with 105k fine-grained caption evaluations over 17.5k audio clips. Our perception–reasoning cascade consolidates descriptions from selected audio-language models into high-quality synthetic references, against which GPT-5-mini evaluates candidate captions to generate scores and explanations as training targets. Fine-tuned on this supervision, Audio-Compass evaluates audio–text alignment directly without references at inference time, supporting both audio-to-text caption evaluation and text-to-audio generation evaluation. Across BRACE, RELATE, PAM, and our new long-form benchmark BRACE-Long, \model shows strong agreement with human judgments, achieving a Spearman correlation of 0.636 on PAM and a temporal-relation correlation of 0.370 on RELATE(OS), compared with up to 0.295 for the evaluated CLAP-based metrics. Further analysis suggests that holistic human judgments are more strongly associated with sound-event information than with acoustic attributes or temporal relations, motivating explicit component-wise evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.