acceptodds
Under review as a conference paper at ICLR 2027

Linear Readability Does Not Guarantee Writability: Evidence from Controlled Counterfactual Recomposition

Abstract

Linear probing reveals information encoded in large language models, but whether probe-derived directions can reliably steer model decisions remains unclear. To investigate this question, we carefully design a controlled experiment on tasks constructed from ALFWorld and WebShop trajectories, whose action lists allow us to define explicit state changes and determine the corresponding answers. In the experiments, we combine the source input's queried state with the target input's decision rule and answer mapping, which we term counterfactual recomposition, to test whether steering supports the intended recomputation rather than merely copying an answer or source plan. Simultaneously, we learn and validate a reference steering subspace to ensure that steering can produce the required counterfactual answer at the selected location. This reference enables a controlled comparison with probe-derived steering in the complement space under fixed intervention conditions. Across four models and both environments, probing accuracy remains near-perfect in the complement space, while steering with direct probe directions, covariance-transformed directions, or combinations of multiple probe directions achieves near-zero counterfactual recomposition accuracy. These results reveal a critical finding: high readability alone does not guarantee writability through probe-derived interventions even with effective steering available, which provides valuable insight for developing reliable steering methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.