acceptodds
Under review as a conference paper at ICLR 2027

Region-Prototype Attention for Calibrated Few-Shot Multi-Label Audio Tagging

Abstract

Few-shot multi-label audio tagging, the task of assigning several co-occurring sound labels to a clip from only a handful of examples per novel class, has been studied only in narrow, non-standardized settings, and we find that standard prototypical models expose a subtle but serious failure: while their ranking of labels is strong, their decisions are badly miscalibrated, with models over-predicting and assigning almost every candidate label to each query. We make three contributions. (i) We establish a few-shot multi-label benchmark on FSD50K and AudioSet-Balanced with a disjoint base/novel class split, a chance floor, and identically-trained baselines in single-class and compositional (LC-Protonets-style) prototypical networks, a frozen AudioSet Transformer, and a post-hoc temperature-scaling calibration baseline. (ii) We show that over-prediction and miscalibration are pervasive across convolutional, compositional, and transformer backbones; a large AudioSet-pretrained Transformer attains high recall only by predicting nearly all labels, with an expected calibration error (ECE) far above that of metric-based models. (iii) We introduce RAFTa, whose region-prototype readout lets class prototypes attend over a query's time–frequency regions before matching. RAFTa is the best-calibrated model in every setting, with the lowest ECE among all baselines and the predicted label-cardinality closest to the truth, while remaining competitive on ranking (mAP), and its calibration edge is not matched by post-hoc temperature scaling of a strong baseline. Code and benchmark splits are released via an anonymized repository.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.