One Concept, Two Facets: Rethinking the Readout-vs-Control Verdict in Large Language Models
Abstract
Linear probes, classifiers fit to a model's intermediate representations, decode high-level concepts from residual-stream activations with near-perfect accuracy, yet a decodable direction does not always steer generation. Prior work attributes this gap mainly to properties of the concept, but cannot explain why the same concept yields opposite verdicts under different extraction methods. We show that the readout/control verdict is jointly set by the concept's structure and the extraction pipeline, and distinguish two poles: a readout direction whose activation traces the surface tokens realizing a concept, and a control variable direction that shapes downstream generation independently of those tokens. To test this claim, we build a concept bank balanced by concept structure and develop a structure-adaptive protocol that derives the steering layer, budget and coherence threshold from the model under test, scoring effects against a norm-matched random-direction null across five independent model families. Concept structure is positively correlated with steerability in all five families on both splits, and the held-out correlation is significant in every family. Holding the concept and data fixed, selecting attention heads by separability recovers the readout facet while an unselected residual difference-of-means recovers the control facet. Alignment audits that read concepts off such directions may therefore miss what the model acts on.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.