acceptodds
Under review as a conference paper at ICLR 2027

Reasoning and learning about injected concepts in language models

Abstract

Language models can, on occasion, correctly answer questions about "injected concepts", i.e., steering vectors added to their activations. This capability, however, remains fragile across models and prompts, leaving open how far it generalizes. To investigate this, we test whether in-context examples help models (i) report properties of the injections and (ii) condition their responses on the identification of specific injected concepts. First, we query about the magnitude of the injection. Second, we ask models to classify whether an injection was applied at an early, middle, or late layer, across different injected concepts. Third, we ask models to gate their output conditioned on the detection of a particular concept, e.g., multiply the answer to an arithmetic problem by if, and only if, an "anger" injection is detected. We test five open-weight models, drawn from three families and two sizes. We find that some models achieve near-perfect performance on magnitude or layer classification and, notably, generalize to unseen magnitudes or layers. Some models also learn to condition their behavior on the identification of specific injected concepts, reaching near-perfect accuracy in selected settings, however with substantial variation across models, tasks, and concepts. These results suggest that some models have useful inductive biases for learning to report nontrivial properties of their activations. This capability motivates new potential paths forward in (automatic, or self-)interpretability research. For example, one can use labeled mechanistic interventions (e.g., steering vectors) to calibrate models to read out properties of their activations, then test whether their reports remain informative when queried without the intervention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.