The Spectral Shortcut: How a Fast Contextual Pathway Suppresses Spectral Learning in Serial Spectral–Spatial Models
Abstract
In this paper we study shortcut learning in serial spectral–spatial models for infrared tissue classification: an encoder processes each pixel's spectrum, and a contextual head reads the encoded neighbourhood. Changing their relative training rates can alter the learned representation while leaving the fitted classifiers similarly context-dependent. In a controlled synthetic experiment, slowing the head by a factor of 256 raises mean linear-probe accuracy from 0.65 to 0.83 at matched training loss, while original-classifier accuracy under contextual reversal changes only from 0.12 to 0.14. Retraining a fresh head for 20,000 steps on decorrelated-context data, with the encoder frozen, yields a mean paired reversal-accuracy gain of 0.17 for slow-head over fast-head encoders, positive in all nine evaluable pairs; one of ten prespecified slow-head runs does not reach the matching loss. The mean recovery advantage persists at all tested combinations of three learning rates and three budgets up to 80,000 steps. A solvable logistic model establishes a sufficient mechanism: contextual readout width and learning rate jointly control a bound on spectral adaptation, with reversal failure under explicit initialization and finite-width conditions. Readout normalization removes the width dependence without necessarily preventing failure. A two-layer ReLU encoder reproduces the accessibility and recovery orderings. Constructed-context experiments with infrared tissue spectra and public hyperspectral spectra reproduce suppression and reliance, but the effects of slowing the head vary across settings. In the high-amplitude tissue construction, blocking neighbour-derived encoder gradients restores strong accessibility and recovery responses to slowing the head. In the tested natural 3×3 neighbourhoods, reliance occurs without a detected encoder-accessibility deficit. Together, these results distinguish contextual reliance from the usefulness of the learned representation for subsequent head retraining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.