Reasoning and learning about injected concepts in language models
Abstract
Language models can, on occasion, correctly answer questions about "injected concepts", i.e., steering vectors added to their activations. This capability, however, remains fragile across models and prompts, leaving open how far it generalizes. To investigate this, we test whether in-context examples help models (i) report properties of the injections and (ii) condition their responses on the identification of specific injected concepts. First, we query about the magnitude of the injection. Second, we ask models to classify whether an injection was applied at an early, middle, or late layer, across different injected concepts. Third, we ask models to gate their output conditioned on the detection of a particular concept, e.g., multiply the answer to an arithmetic problem by if, and only if, an "anger" injection is detected. We test five open-weight models, drawn from three families and two sizes. We find that some models achieve near-perfect performance on magnitude or layer classification and, notably, generalize to unseen magnitudes or layers. Some models also learn to condition their behavior on the identification of specific injected concepts, reaching near-perfect accuracy in selected settings, however with substantial variation across models, tasks, and concepts. These results suggest that some models have useful inductive biases for learning to report nontrivial properties of their activations. This capability motivates new potential paths forward in (automatic, or self-)interpretability research. For example, one can use labeled mechanistic interventions (e.g., steering vectors) to calibrate models to read out properties of their activations, then test whether their reports remain informative when queried without the intervention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.