Shortcut Lock-In: The Dynamics of Learning and Escaping Spurious Features
Abstract
Neural networks can improve at the core task while continuing to follow a spurious cue on conflicting inputs. When does further training allow useful task information to take control? We investigate this gap between capability and decision control through an evidence-transport account that separates the supply of corrective examples from the decision progress induced by their updates. A tractable two-route model shows how continued core learning during shortcut dominance leads to eventual escape under persistent conflict evidence, with distinct capture and escape timescales. During recovery, ViT-Tiny preserves coarse semantics while its sample geometry changes more than ResNet-18’s. Early measurements also carry timing information: on four prospectively held-out ViT-Tiny runs, first-ten-update estimates predict escape onset with 13.33% median absolute percentage error. Visual and multimodal experiments reproduce capture and recovery, and a five-seed GQA study extends the analysis to natural image–question pairs. These results distinguish the acquisition of useful task information from the training dynamics that make it decisive.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.