acceptodds
Under review as a conference paper at ICLR 2027

Are Frequency-Balanced Datasets Optimal for Equitable Class Performance?

Abstract

Resampling is a widely used approach for mitigating class-wise performance disparities in frequency-imbalanced datasets, typically by reallocating training data according to class frequency. However, this framing treats representation as the primary basis for deciding how much data each class should receive, while classes may also differ substantially in learning difficulty. Hardness-aware resampling and reweighting have shown that allocating more training resources to harder classes can improve class-wise fairness, but existing evidence comes predominantly from frequency-imbalanced settings, where the effects of correcting representation and accounting for hardness are difficult to disentangle It therefore remains **unclear whether hardness provides a useful basis for data allocation in frequency *balanced* settings**. We investigate this question by deliberately allocating more training data to harder classes. We find that this reallocation reduces class-wise recall disparities while maintaining comparable overall performance. We further show that the effectiveness of oversampling depends on the quality of the added samples, with hardness emerging as a strong estimate for downstream task utility. These results demonstrate that **resampling can be useful even when there is no representation imbalance to correct**, and motivate greater attention to class-wise fairness in frequency-balanced settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.