Holding the Answer Fixed: What Hidden-State Probes and Interventions Establish About Correctness
Abstract
Hidden-state probes are often interpreted as measuring whether a model's answer is correct. But correctness is relational: the same answer can be right for one task and wrong for another, while standard evaluations usually allow both the task and the answer to vary. We study what probe success establishes when answer identity is held fixed. We construct paired task variants in which the answer is identical but its checker outcome flips. Across six 7B–32B code and tool-use models, newly fitted probes almost perfectly distinguish these fixed-answer task conditions. Yet this strong decodability does not imply that an existing correctness score is reusable: probes learned from freely generated answers transfer inconsistently to the same candidate states and can even reverse their ordering under the tested shift. We uncover a second separation within a single frozen score. Two probes correctly order 93% and 95% of fixed-answer pairs, while incorrect answers still outrank correct answers from other problems in about one third of cross-problem comparisons; generation-matched mazes exhibit the same gap. Input controls show that the surviving signal can arise from different task-dependent sources. Finally, late whole-state patches reproduce donor outputs whether those outputs are correct or incorrect. Together, these results show that decodability, score reuse, fixed-answer responsiveness, cross-problem ranking, and output restoration are distinct properties. Correctness probes should therefore be evaluated against the comparison required by their intended use, rather than by pooled discrimination alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.