acceptodds
Under review as a conference paper at ICLR 2027

Same Data, Different Target: How the Prediction Target Can Separate Probing from Steering

Abstract

A linear probe can decode a fact while steering along it barely changes the model's output, and such a failure is easily read as the model not using the fact. We ask whether it can instead reflect what the model's readout, its output layer, was trained to predict. We train small blocks-world transformers that share data, initialisation and batch order and differ only in whether the readout predicts the current, next or previous state. Under the next- or previous-state target, clear(x) stays decodable, but its difference-in-means direction aligns more with the readout row of holding(x) than with its own. In our random play, holding(x) is possible one step away only if clear(x) holds now, and a model fitted to the labels alone predicts where the direction goes. How much steering along it changes the clear(x) output depends on its component on that output's row relative to the outputs' logit margins: it is partial where next-state outputs are uncertain and, in distribution at our primary strength, near zero where previous-state outputs are confident, under 3% of what the row itself achieves. The alignment replicates on new seeds and on one encoding whose action counts do not fix the state, and in Othello the target sets the sign of a probe direction's alignment with its own row. A steering failure can therefore reflect the readout's target rather than the representation; where the readout is hidden, a probe fitted to the model's own outputs diagnoses the alignment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.