The Role of Data-Reuse on Shortcut Learning with Linear Networks
Abstract
Many dynamical theories of shortcut learning use fresh samples or population gradients, whereas most practical training repeatedly uses a finite dataset. We study full-batch gradient descent with square loss in a synthetic linear model of spurious correlations. Our analysis uses a non-rigorous dynamical mean-field calculation from statistical physics. The resulting equations separate shortcut and core learning from fitting directions specific to the training dataset. A perturbative analysis predicts a reuse crossover on a timescale proportional to , where is dataset size and is input dimension. Without regularisation, the interpolation peak at can degrade test performance without impairing core or shortcut acquisition. For , overlap between the features can amplify the shortcut coefficient. Ridge suppresses this peak but can increase the ratio of shortcut to core coefficients. Simulations validate the theory, while image benchmarks provide qualitative evidence beyond the solvable setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.