acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Class Reweighting through Temporal Evidence for Long-Tailed Recognition

Abstract

Long-tailed audio recognition typically measures class imbalance by training sample counts per class. This view overlooks how useful information is distributed within each sample: the labeled sound may occupy only a small portion, while persistent or repeated patterns make other portions redundant. Classes with identical sample counts therefore differ in the amount of irrelevant or repeated content. To account for recording content in class reweighting, we propose Constrained Effective Temporal Evidence Learning (CETE) for long-tailed audio recognition. CETE introduces a temporal head to identify temporal features that support the class decision. The head is trained with a loss that combines classification constraints with a penalty on the total temporal weight. This loss encourages correct classification after feature weighting while discouraging large weights on unnecessary background or repeated features. A teacher model selects correctly and confidently classified examples for learning temporal weights and computing class statistics. CETE aggregates their temporal weights within each class and combines the resulting statistics with sample counts to determine class loss weights, reducing bias toward common classes. Experiments on VGG-Sound-LT, iNatSound, and BirdCLEF demonstrate CETE's consistent superiority over the baselines and its robustness across controlled imbalance ratios and naturally long-tailed data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.