acceptodds
Under review as a conference paper at ICLR 2027

The Hiding Problem: When Unread Features Do and Do Not Drive Harmful Adaptation

Abstract

When a classifier adapts to new data without labels and its loss rises, the loss alone does not say why: the model may have acquired a harmful dependence on a component of its representation that it did not use before, or it may have become more confident in the mistakes it already made. We ask which of these happens when a later update reaches components that the current predictor does not read (the Hiding Problem). An exact binary construction (Theorem 1) shows that the first is possible: two predictors with identical calibrated predictions and zero population IRMv1 penalty take the same Euclidean entropy step, and the one whose unread coordinate is stretched climbs from 75% to 90% accuracy on the adaptation law and falls to 10% under a reversed nuisance association, while the other stays at 75% under both. Deliberate amplification produces the same behavior in synthetic and frozen-feature Camelyon17 representations. On six native ResNet-50 checkpoints for a naturally shifted hospital, three controls point to the second explanation instead: a label-free affine map of the initial score closely reproduces over 99% of the head-only loss increase, a selected unread direction keeps no mean advantage over random deletions once endpoint movement is matched, and the same parameters trained with 320 labels raise the ranking in every checkpoint, whereas label-free updates that reach the representation lower it in the three checkpoints that missed the most positives (a post hoc split). A preregistered replication on Waterbirds-95 with six fresh models reproduces the head-only account and the sign pattern of the representation arms; the widespread one-class collapse at a fixed learning rate did not replicate. Identical predictions do not imply identical responses to learning. An update can recruit an unused feature, increase confidence in existing mistakes, or appear less harmful simply because it changes the predictions less. Our results make these distinctions consequential: evaluating continued learning requires testing what an update changes and establishing which changes account for its harm; the checks here begin that test and lay the groundwork for safeguards that can be tested the same way.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.