AudioCaliper: A Controlled Probe of Audio-LLMs as Speech Quality Judges
Abstract
Audio large language models (audio-LLMs) now support reasoning beyond simple classification. They are increasingly used to train other models. They curate speech data, score text-to-speech outputs, gate deployments, and post-train generators as reward models. However, whether they can be trusted in these roles is rarely tested. In this paper, we present AudioCaliper, an open-source, reusable framework for analyzing how an audio-LLM behaves as a quality evaluator. For each clip, we vary one factor at a time while keeping the rest fixed, either the acoustic degradation, the presence of speech, a quality claim in the text prompt or the answer position. This setup allows us to attribute changes in the output to the manipulated factor. We apply AudioCaliper to LibriSpeech and VCTK using Qwen2-Audio, Qwen2.5-Omni, SQ-LLM, SpeechJudge-GRM, SALMONN-SQA and QualiSpeech. Our results reveal each judge's sensitivity to the audio input and text prompt. Fine-tuned judges reach an acoustic sensitivity of 0.87, while zero-shot judges near zero. However, fine-tuning does not eliminate other failures. Three out of six judges, two of them fine-tuned, cannot separate speech from silence, and two score silence higher than real speech. All six judges remain sensitive to quality claims in text prompts. Describing a clip as high quality rather than low quality moves a fine-tuned judge's score by 31% of its rating scale and a zero-shot judge's by 86%. In pairwise comparisons, five out of six judges rely on the position of the clip rather than the input audio. Our findings call for evaluating a judge on the decisions it will support, whether scoring, ranking or rewarding, before its scores are used to train or evaluate other models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.