acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Inference-Time Safety Alignment via Attention Routing

Abstract

Despite extensive safety alignment, large language models remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense, but conditional steering still needs a reliable signal for deciding to intervene. Existing gates often read residual stream projections, yet jailbreak transformations can move these projections toward the benign regime. We instead examine attention routing at the assistant header, where refusal-related activation concentrates. A jailbreak alters how a small set of heads route the preceding context at these tokens, whereas benign and directly harmful prompts exhibit similar routing patterns. We aggregate the signed deviations of these heads from their benign routing pattern into the (RDI), which captures jailbreak-induced routing changes rather than the surface form of the rewrite or the harmfulness of the underlying request. We build (outing-ndexed afety lignment) on this routing measurement. In , a jailbreak is detected once that deviation passes a threshold set on benign prompts. In , the excess above that threshold sets the strength of a bounded refusal update, which is applied in a second prefill pass only when the gate is triggered. Across three open-weight aligned chat models and seven jailbreak families, RISA-D achieves the highest average detection accuracy on all three models, while RISA-S raises the global mean defense success rate from 40.0% to 95.0%. The routing signal also remains effective when individual attack families are excluded from calibration. These results suggest that attention routing offers a new perspective on safety alignment at inference time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.