acceptodds
Under review as a conference paper at ICLR 2027

Semantic State Atlas: State-Conditioned Coordinates for Activation Intervention

Abstract

Activation steering can alter model behavior without updating weights, but an edit calibrated at one hidden state need not retain the same effect at another state or representation stream. Local intervention directions can weaken, affect a different behavior, or alter unintended predictions as the anchor changes. We study this problem as semantic persistence. Semantic State Atlas represents an intervention with anchor-specific local coordinates and explicitly transfers those coordinates between charts. Each chart selects directions that respond to the requested target while limiting changes to held-out non-target predictions. Transition operators are trained so that a transported edit reproduces the source edit's measured target responses at the destination, with consistency imposed only across semantically equivalent routes. A denoising-based validity surrogate further regularizes edited states toward locally supported regions. On controlled semantic-binding tasks, Chart and Atlas perform almost identically at one hop (91.5% versus 91.6%), but the gap grows to 73.0% versus 85.1% at three hops and 53.2% versus 71.8% at six hops. A matched state-conditioned MLP reaches 76.7% and 59.3%, respectively, indicating that state conditioning alone does not account for the full gain. With paired stream maps fixed, mean cross-stream success rises from 63.8% to 67.5%. Single-step benchmark gains are smaller, while nested procedural edits and added inference cost remain limitations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.