acceptodds
Under review as a conference paper at ICLR 2027

Expressed Interpretation Vectors in LLMs

Abstract

Large language models can often glean the intent of a user's question despite typographical errors. However, past work has shown that sufficiently corrupted prompts can induce a lack of attempted interpretation. In this work, we explore whether this behavior is linearly represented in the activation space of models. To this end, we construct a contrastive dataset of prompts by corrupting questions until models transition from interpretation to non-interpretation, extract candidate directions via difference-in-means, and perform selection by evaluating the behavioral effects of each one during steering. We find that across nine instruction-tuned models, projecting the direction out of the residual stream increases attempted interpretation by 39 to 70 percentage points on held-out corrupted TriviaQA questions, while adding it induces non-interpretation. The effects of these interventions transfer to SimpleQA, and in the largest model of each family, to word-level omission and OCR corruptions without re-extracting the direction. Notably, across six utility benchmarks, direction removal only changes scores by a mean absolute change of 0.74 percentage points, showing that general capabilities are left largely intact. We also demonstrate that the direction is separable from the refusal direction. Ultimately, our results suggest that whether a model attempts to interpret a degraded prompt is mediated by a low-rank subspace: several directions within it restore interpretation, but only directions built from our interpretation contrast reliably suppress it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.