acceptodds
Under review as a conference paper at ICLR 2027

Can Context Arbitrarily Reorient Linear Representations in Large Language Models?

Abstract

The so-called concept vectors in the activation space of language models are widely used to interpret and steer high-level behavioral traits represented in model outputs. These directions are typically extracted from a given pair of contrastive sample sets and treated as fixed for the model at hand thereafter. Complementary to the emerging literature on the effects of input context on linear representations, we systematically analyze how strongly context alone can change these directions while holding the model, layer, and the sample sets fixed. We find that optimizing a few shared input prefix tokens can reorient the extracted concept vector to a wide range of arbitrarily chosen directions, including the direction opposite to the original one. Further analysis reveals that reorientation of the extracted contrast vector is achieved by prefix-induced displacement in the underlying activation geometry. Preliminary observations suggest that a related geometric signature may also accompany successful adversarial prompts such as Greedy Coordinate Descent (GCG). Incorporating this geometric objective into GCG indeed yields modest gains in attack success on some models. Furthermore, we take a step toward reducing multi-concept interference under activation steering by optimizing the prefix to make the extracted contrastive vectors for different concepts orthogonal. Overall, our findings suggest that reliable interpretability should treat linear representations as context-conditioned objects rather than as specific only to the model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.