Lagged Coupling: Decodability and Probe-Direction Interventions Across Training
Abstract
Information that is linearly decodable from a language model need not be accessible through an intervention along the decoding direction. We track this distinction through training across six Pythia models (160M–12B parameters), eight checkpoints, and four synthetic task families. At each model–checkpoint–task combination, we measure internal probe AUROC, output-level answer-score AUROC, and the effect of a single-site additive edit along the probe direction relative to norm- and site-matched random directions. The three measurements dissociate. Internal decodability is near ceiling (probe AUROC ) from the first measured checkpoint; output-level discrimination, tabulated for the four largest models, rises by up to 0.43 AUROC over training; and 43 of 48 intervention summaries fall within the random-direction band , with no sustained positive effect and a consistently negative alignment at the two earliest checkpoints. Over the same period, the RMS activation along the probe direction grows by up to 56.8-fold without a corresponding intervention benefit. We call this dissociation lagged coupling. A first-order analysis explains why it is expected: the standardized intervention effect equals times the cosine between the probe direction and the model's mean readout gradient, independently of the edit norm, and near-perfect decodability places no lower bound on that cosine when the representation is anisotropic. Probe accuracy is therefore not a proxy for steerability, and decodability and intervention efficacy should be evaluated separately throughout training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.