acceptodds
Under review as a conference paper at ICLR 2027

Decodable Is Not Necessarily Used: Probe Regularization Controls the Causal Relevance of Belief-State Representations in Transformers

Abstract

Linear probes are the default tool for reading concepts out of neural representations, and a high decoding accuracy is routinely treated as evidence that the model uses the decoded feature. We show, in a setting with an exact analytic ground truth, that this inference can fail almost completely. We train transformers on hidden-Markov processes (Mess3 and the Even Process) whose optimal Bayesian belief states are known in closed form, and we measure causal use by counterfactually patching a belief representation and asking whether the model’s next-token distribution moves to the belief-implied Bayes-optimal target (β=1 means it moves exactly as an ideal belief-user would). We find that belief is linearly decodable at R2 ≈ 0.99(Mess3) from a continuum of residual-stream subspaces spanning nearly 90◦, yet the causal efficacy of these subspaces varies from β ≈0 to β ≈0.85. Decoding accuracy is nearly invariant across this fan; causal efficacy is not. Unregularized least-squares probing lands on a subspace that is near-orthogonal to the functional directions and only weakly read by the model, whereas ridge regularization rotates the recovered subspace toward the high-variance functional directions that are independently identified by both PCA and Distributed Alignment Search. We rule out magnitude/LayerNorm, token-identity, variance-destruction, and structural-register confounds, and verify the effect is belief-specific, history-dependent, present at every layer, and consistent across seeds and three generative processes. Our central methodological result is that probe decoding accuracy carries almost no information about causal relevance: two subspaces with identical R2 can differ in causal effect by two orders of magnitude, and which one a probe recovers is silently determined by its regularization. We argue that probe-based claims of use require explicit causal validation and reporting of regularization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.