Farsighted Correction of Shortcut Learning: From Dynamic Starvation to Confidence-Gap Optimization
Abstract
Deep neural networks frequently rely on easily accessible spurious correlations, driving a phenomenon known as shortcut learning. Remarkably, this reliance remains invisible on standard loss curves, which decrease smoothly as models build brittle representations. In this work, we uncover the mathematical root of this invisibility by analyzing optimization dynamics via the empirical Neural Tangent Kernel (eNTK) along the training trajectory. We establish the non-asymptotic **Dynamic Starvation Theorem**, proving that a dominant shortcut eigenvalue restricts causal parameter updates to a bounded regime in the worst case of unbounded shortcut dominance. We further characterize how this starvation mechanism manifests under different loss functions: under **squared loss**, joint and ablated training yield identical predictions in output space, trapping the causal subnetwork in a degenerate null space as loss vanishes; under **cross-entropy loss**, self-correction is merely logarithmic. To resolve these failure modes, we propose **Farsighted Signal Optimization** (FSO), which reinjects counterfactual causal gradients into the training objective. We extend this to model-agnostic settings via GAP-FSO, proving that inter-group confidence gaps serve as exact proxies for shortcut parameters without explicit architectural splits. Unlike prior theoretical treatments of shortcut learning, which validate their predictions on simple architectures, we test GAP-FSO on ResNet-50 and DistilBERT across six vision and language benchmarks. Results are competitive with established baselines overall and clearly outperform them on the NLP benchmark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.