Expressed Interpretation Vectors in LLMs
Abstract
Large language models can often glean the intent of a user's question despite typographical errors. However, past work has shown that sufficiently corrupted prompts can induce a lack of attempted interpretation. In this work, we explore whether this behavior is linearly represented in the activation space of models. To this end, we construct a contrastive dataset of prompts by corrupting questions until models transition from interpretation to non-interpretation, extract candidate directions via difference-in-means, and perform selection by evaluating the behavioral effects of each one during steering. We find that across nine instruction-tuned models, projecting the direction out of the residual stream increases attempted interpretation by 39 to 70 percentage points on held-out corrupted TriviaQA questions, while adding it induces non-interpretation. The effects of these interventions transfer to SimpleQA, and in the largest model of each family, to word-level omission and OCR corruptions without re-extracting the direction. Notably, across six utility benchmarks, direction removal only changes scores by a mean absolute change of 0.74 percentage points, showing that general capabilities are left largely intact. We also demonstrate that the direction is separable from the refusal direction. Ultimately, our results suggest that whether a model attempts to interpret a degraded prompt is mediated by a low-rank subspace: several directions within it restore interpretation, but only directions built from our interpretation contrast reliably suppress it.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.