Allocating Weight Precision for Deep Imbalanced Regression with Edge-Case Annealing
Abstract
Deep Imbalanced Regression (DIR) addresses the mismatch between naturally imbalanced training distributions and the more balanced performance desired across the target range. Methods that resolve the underlying capacity conflict, in which parameters shared across the target range receive updates dominated by majority samples, currently do so by restructuring the model and therefore sacrifice orthogonality. We introduce Edge-Case Annealing (ECA), a training-schedule method that leaves the loss and the architecture untouched. ECA couples a data exposure curriculum one to one to the learning rate schedule: a ceiling on the target histogram tracks the learning rate, so that the effective training distribution reaches balance exactly as the step size vanishes. The design follows a minimum description length view of capacity: under the imbalanced training distribution, fine-grained structure for rare targets does not justify the weight precision it requires, whereas under the balanced evaluation distribution it does, and it can be retained once the learning rate, and with it the gradient noise, has decayed. On AgeDB-DIR, IMDB-WIKI-DIR, and the MuViS-DIR time series benchmark, ECA is the strongest orthogonal addition, improves the tail without the majority tax of static rebalancing, and composes with existing methods up to a new state of the art in the medium and few-shot regions. Under increasing weight noise, the annealing gains are erased in rare target regions, consistent with being stored as small, precision-fragile adjustments and supporting the proposed capacity view.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.