acceptodds
Under review as a conference paper at ICLR 2027

Learning Environment Distributions for Robust Image Classification

Abstract

Machine learning models often exploit spurious correlations, such as using background to recognize a bird in Waterbirds or the color of a digit to classify it in Colored MNIST. One approach to reducing this reliance is to train across multiple environments: data distributions that differ, for example, in their color-label correlation or in how training examples are grouped. However, the choice of these distributions and their contribution to training can strongly affect the resulting predictor. In this paper, we introduce Universal Adaptive Environment Discovery (UAED), a framework that jointly learns a predictor and a distribution over a specified family of environment constructions, regularized by a KL penalty toward a reference. For transformation-based training, this distribution determines how different transformations contribute to the objective; for group-based training, it weights alternative groupings of examples, which we construct from a reference predictor's prediction difficulty. Our analysis characterizes the resulting policy preferences: learned cooperatively with the predictor, the distribution tilts toward environments the predictor already handles well and collapses onto a shortcut environment below a computable KL threshold, whereas reversing the sign of its update yields a robust counterpart whose objective upper-bounds the expected training score over a KL ball of environment distributions. Experiments distinguish the effects of broader coverage from those of learning its weighting. Training over the family improves worst-case accuracy over restricted environments on Colored MNIST and worst-group accuracy over ERM on Waterbirds with ResNet-50 without group annotations, while with ViT-B/16 the grouping objective destabilizes fine-tuning at the shared learning rate. Learned policies move as predicted, away from the annotated grouping when cooperative and toward it when robust, and weak KL penalties reinforce color shortcuts; on the official GroupDRO code, where the grouping family reaches 81.8% worst-group accuracy without group annotations, cooperative learning lowers accuracy in every seed and the robust policy removes this loss, matching but not exceeding a fixed weighting. These findings identify both the potential and the limitations of jointly optimizing predictors and environment distributions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.