acceptodds
Under review as a conference paper at ICLR 2027

Inference-Time Attention Manipulation

Abstract

Open-weight language models expose their internal computation at inference time, creating a safety surface that is not captured by robustness to adversarial prompts alone. We introduce a novel white-box threat model based on **region-level attention manipulation**, which tests refusal behavior by steering or cutting attention between semantic regions of a chat. Under this threat model, each manipulation modifies inference-time attention while holding model parameters fixed, requiring neither training nor adversarial prompt optimization. *Attention cutting* from the decision region to the harmful request can bypass refusal in several models without modifying the prompt or forcing a prefill, while some models resist this manipulation. For models that resist, attention manipulation during continuation can still bypass refusal, suggesting that attention dynamically maintains refusal beyond the response opening. We further show that *attention steering* toward a supplied harmful prefill increases attack success across most evaluated model families, although stronger interventions do not consistently produce larger effects. Together, these results show that refusal depends not only on what appears in the context but also on where and when the model allocates attention during generation. Safety evaluations of open-weight models should therefore test attention-level interventions both at the decision region and during continuation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.