acceptodds
Under review as a conference paper at ICLR 2027

When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation

Abstract

A common response to spurious correlations is to weaken them, for instance by balancing the training data so that the shortcut feature no longer predicts the label. We find that in small transformers this can make the robust rule harder to learn. In our tasks, a shortcut feature agrees with the label on a fraction of training examples and contradicts it on a held-out adversarial split. When the label is the sum parity of an integer sequence and the shortcut is the parity of its largest element, the fraction of two-layer seeds that reach adversarial accuracy grows from 0% at to 53% at , whereas for one-layer models it drops from 33% to 0%. The benefit depends on how far is from chance rather than on the direction in which it deviates. On a ternary task, two-layer models generalize in 80–93% of seeds both when the shortcut is correlated and when it is anti-correlated with the label, and in no seeds at chance. In that setting they memorize the training set, a failure that disappears with 16 more data. All models learn the shortcut-only predictor first, and in the runs that succeed the robust rule emerges while attention to the shortcut token remains high. Finally, models trained briefly with an informative shortcut and then moved to balanced data still generalize in most seeds, which suggests that the shortcut serves as an optimization scaffold rather than working by amplifying the gradient of the misclassified minority.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.