FullSongEval: Auditing Text–Label Associations Before Music Evaluation
Abstract
Public music datasets pair text with labels such as genre. Before using these labels in evaluation, researchers may ask whether the text predicts them. Yet a scorer can remain above label prevalence after complete label sets are reassigned among records. FullSongEval audits a candidate rule that accepts text–label pairings when a paired fold interval for a scorer’s average-precision margin over prevalence has a positive lower endpoint. We test original-pair acceptance and scorer sensitivity to an injected label-linked cue. We then refit the scorer after two fixed label-set reassignments per dataset and ask whether the rule rejects them. On MusicCaps and MTG-Jamendo, the scorer detects the cue, but the rule accepts both reassignments. On the same MTG tracks, audio-feature genre classifiers trained with original tags predict independently collected consensus genres more accurately than classifiers trained with reassigned tags: mean average precision is 0.7927 versus 0.4365 and 0.4304. A diagnostic finite-sample random-label reference accounts for tied scores. In a separate 8,000-record Free Music Archive check fixed before source inspection, this reference separates original pairings from two reassignment controls, while the prevalence rule accepts all three. Exceeding prevalence is insufficient to approve these text–label pairings for later music evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.