When Are Spectral Modes Causal? Identification Before Intervention in Learning Dynamics
Abstract
Spectral transitions (outliers in the loss Hessian or in parameter-update matrices) are increasingly read as mechanistic signatures of feature learning and grokking, but such readings are observational: whether an emerging spectral mode controls feature acquisition or merely marks it is rarely tested. We give a preregistered, intervention-based answer. We argue a spectral direction is a candidate causal object only after passing an explicit identification stage (prospective detection, external separation, projector stability, and a quantitative feature-specificity gate), and that whether it is manipulable then depends on the operator in which curvature is measured. Our central finding is a clean dissociation: the causal object exists but lives in different spaces for different problems. In an analytic teacher-student model, prospectively suppressing the detected full-Hessian outlier slows feature-alignment growth by versus a norm-matched-random floor ( seeds, ), with a monotone dose-response, an energy-matched control isolating direction from removed update amount ( effect), a positive-curvature-branch localization, and a theory-predicted / dissociation with placebo interactions equivalent to null. In deep networks the parameter-space operator no longer suffices: across four grokking modular-addition transformers the prominent, pre-generalization, feature-aligned full-Hessian mode is not an identifiable target (no rank in - jointly clears separation, stability, and specificity, under both the full Hessian and a validated true Gauss-Newton operator), and a planted-object calibration (recovery once a spike clears the bulk by , zero false positives) shows this is a genuine identification gap, not an insensitive assay. Yet the feature is not unidentifiable, only misplaced: a function-space input-sensitivity operator recovers a deep-net mode that passes every gate and is causally manipulable, its targeted suppression delaying generalization (median steps) far beyond norm-matched-random, energy-matched, and wrong-bits controls (censoring-robust rank test against each, seeds). A planted-object calibration confirms this operator is not blind on transformers; it simply finds no compact object there under any tested operator, including the symmetry-adapted per-frequency operator the task most invites, so the feature is distributed rather than a detector artifact. Identification before intervention thus does more than reject over-readings of spectral edges: it tells you which operator exposes the causal object, and flags when no tested operator yields a compact one. All claims are preregistered with hashed-before-results protocols.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.