acceptodds
Under review as a conference paper at ICLR 2027

When Linear Probes Go Blind: The Order Spectrum of Learned Representations

Abstract

Linear probes are widely used to assess properties encoded in neural network representations. However, poor linear recovery cannot distinguish absent information from information accessible through nonlinear readouts, while commonly used MLP heads do not characterize how information is distributed across orders. We formalize this distinction by introducing the order spectrum, a Hermite decomposition of the target regression function that, under Gaussian representations, resolves predictable variance by polynomial degree. Specifically, we show that optimal affine probes measure the first-degree energy of the spectrum, while increments between optimal nested polynomial probes recover successive energies. For the off-Gaussian case, we bound the discrepancy in terms of the Hermite-feature Gram deviation and a target-dependent approximation residual, and validate the bound in controlled settings. Leveraging this framework, we show that quadratic probes recover substantial board-color information in OthelloGPT and pendulum energy in a JEPA trained without energy labels, where linear probes fail. Further, by analyzing grokking in modular addition, we find that partial linear recovery can coexist with substantial quadratic accessibility, while quantitative Hermite-energy interpretations require explicit distributional and approximation conditions. More broadly, the framework provides a degree-resolved account of how target information becomes accessible in learned representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.