How Rare Counterexamples Change What Models Learn from the Majority
Abstract
Models often rely on shortcuts that predict most training examples correctly but fail on rare counterexamples. We study why the same counterexamples can correct shortcut reliance at different rates depending on the remaining training data. In a two-feature model trained with logistic loss, we vary how robust and shortcut feature values are paired within majority examples while holding the counterexamples, feature marginals, and covariance fixed. We identify an indirect mechanism: counterexample updates change majority-example margins, reactivating gradients that both promote robust-feature learning and reinforce the shortcut. The balance between these effects evolves as different majority examples come to dominate learning. Our analysis links these shifting gradient contributions to feature growth and the time to first correct classification, showing that matched marginals and covariance do not determine correction times and that approximations based on the initially hardest examples can miss the subsequent dynamics. A controlled intervention that removes only majority-example updates to the robust feature reduces the ratio of mean correction times between two pairing conditions from 2.17 to 1.02, supporting the role of this indirect pathway. Experiments with small tanh networks on handwritten digits with synthetic backgrounds further show that feature pairing affects worst-group accuracy even when parameter-update norms are matched. Together, these results show how counterexamples reshape subsequent learning from majority examples, revealing a mechanism for shortcut correction that feature marginals and covariance alone do not capture.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.