AUDITING MUSIC EVALUATORS ACROSS ROLES: FROM CONTROL VALIDITY TO TEACHING UTILITY
Abstract
Music evaluators assess attribute control, predict listener preferences, and provide training supervision. Evidence for one use does not by itself justify another. We develop a framework that connects evaluator selection and improvement through role-specific comparisons of control validity, direct preference quality, and teaching utility. A log-loss counterexample shows that better direct preference predictions need not yield a more beneficial one-step update to a constrained student, even with common examples and soft targets. Control evaluations show how label priors and text truncation affect measured agreement, and use waveform replacement to test whether judgments depend on the audio. For musicality preferences, complementary assessment scores improve direct prediction. Separate, fully fold-isolated teacher–student comparisons identify modest reductions in student log loss relative to matched human-only training; some gains persist after temperature calibration of both supervised and human-only students. On common human targets, comparisons among supervised students inform teacher selection, while comparisons with matched human-only training test whether supervision helps. Human-feedback adaptation lowers the log loss of MuQ-based preference scorers on held-out content relative to continued training on the original labels and temperature calibration, closing the evaluation–improvement loop for direct preference quality. The framework grounds evaluator selection and improvement in evidence matched to the intended role and, for supervision, to the student and learning procedure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.