acceptodds
Under review as a conference paper at ICLR 2027

AuditEar: Expert Audio Quality Assessment with Large Audio Language Models

Abstract

Existing speech-quality assessment systems typically reduce a recording to a scalar score, offering little insight into what is wrong, how severe the problem is, or how it should be fixed. We introduce ProQA, a dataset of 634 speech recordings assessed by professional audio engineers across 18 production-quality axes. Its 11,412 axis-level annotations pair ordinal severity ratings with issue tags, evidence-grounded assessments, and recommended fixes. To expand this costly expert supervision, we apply controlled augmentations across ten axes and compound their severity increments onto the original expert ratings, producing 15,524 audio variants with known quality changes. We use this corpus to train AuditEar, an audio-language model that diagnoses quality issues and recommends corrective actions. Our three-stage training paradigm combines Supervised Fine-Tuning, On-Policy Self-Distillation from privileged expert context, and GRPO with Verifiable Rewards. On a held-out test set, AuditEar achieves relative improvements of 41.2% in per-axis severity accuracy, 64.1% in pairwise accuracy, and 86.4% in holistic fix-F1 over its zero-shot backbone. It also outperforms Audio-Flamingo-Next, Kimi-Audio-7B-Instruct, and MiDashengLM-7B across all evaluated tasks, with gains of up to 2× over the strongest of these baselines. On external benchmarks, AuditEar achieves 3×–45× the correlation of published zero-shot audio-language baselines on QualiSpeech and reaches 74.3% preference accuracy on SpeechEval, compared with 70.4% for audio-language models evaluated zero-shot. These results show that expert, per-axis supervision can turn audio-quality assessment into a grounded and actionable diagnosis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.