acceptodds
Under review as a conference paper at ICLR 2027

AudioCaliper: A Controlled Probe of Audio-LLMs as Speech Quality Judges

Abstract

Audio large language models (audio-LLMs) now support reasoning beyond simple classification. They are increasingly used to train other models. They curate speech data, score text-to-speech outputs, gate deployments, and post-train generators as reward models. However, whether they can be trusted in these roles is rarely tested. In this paper, we present AudioCaliper, an open-source, reusable framework for analyzing how an audio-LLM behaves as a quality evaluator. For each clip, we vary one factor at a time while keeping the rest fixed, either the acoustic degradation, the presence of speech, a quality claim in the text prompt or the answer position. This setup allows us to attribute changes in the output to the manipulated factor. We apply AudioCaliper to LibriSpeech and VCTK using Qwen2-Audio, Qwen2.5-Omni, SQ-LLM, SpeechJudge-GRM, SALMONN-SQA and QualiSpeech. Our results reveal each judge's sensitivity to the audio input and text prompt. Fine-tuned judges reach an acoustic sensitivity of 0.87, while zero-shot judges near zero. However, fine-tuning does not eliminate other failures. Three out of six judges, two of them fine-tuned, cannot separate speech from silence, and two score silence higher than real speech. All six judges remain sensitive to quality claims in text prompts. Describing a clip as high quality rather than low quality moves a fine-tuned judge's score by 31% of its rating scale and a zero-shot judge's by 86%. In pairwise comparisons, five out of six judges rely on the position of the clip rather than the input audio. Our findings call for evaluating a judge on the decisions it will support, whether scoring, ranking or rewarding, before its scores are used to train or evaluate other models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.