acceptodds
Under review as a conference paper at ICLR 2027

Concept Control in Language Models via Activation Straightening and Clamping

Abstract

Representation intervention aims to control concepts in LLMs by intervening on internal activations. While promising, existing methods often compromise general model utility or fail to generalize under distribution shifts. We argue that these failures of faithful control stem from two structural obstacles in latent space. First, concepts are redundantly distributed across layers. Sparse interventions allow the model to bypass the intervention (the "hydra effect"), whereas naively intervening at many layers can over-correct activations and degrade utility. Second, the activation patterns that encode a concept can vary with context, so a simple linear subspace often fails to capture the geometry of the concept and can perturb unrelated capabilities, especially on out-of-distribution prompts. We propose Straighten-and-Clamp (SAC). SAC learns an invertible straightening map, which concentrates concept-related variation, followed by a projection clamp to enforce a target concept code in the straightened coordinates. Because the clamp is idempotent at each site, it becomes inactive once the target code is already satisfied, enabling stable multi-layer deployment. We train SAC in two stages: first, we learn a straightening map amortized across concepts; second, we learn lightweight per-concept clamps via pairwise preference optimization. Experiments on concept induction and suppression of harmful concepts suggest that SAC improves the control–utility trade-off and robustness to distribution shifts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.