Safety Fragility Is a Routing Problem: The Attention Bottleneck Between Harm Detection and Refusal
Abstract
Aligned language models refuse harmful requests most of the time, and that refusal is remarkably fragile. Two lines of research explain the fragility in ways that do not fit together: one localizes safety in a single residual-stream direction or a handful of attention heads, the other finds it confined to a narrow basin in weight space that fine-tuning drifts out of. We show that they describe different components of one mechanism. To unify these findings, we decompose the safety into three components: detection of harm, routing of that signal through attention, and emission of a refusal. Both detection and action are stable under weight perturbation, and what breaks is the routing: a small, non-redundant set of attention heads, which we call the gate, connecting detection to action. Standard feature attribution, the method the field uses to identify which heads matter for a behavior, ranks this set incorrectly in all three model families we test, because it scores how much a head writes toward refusal rather than whether the output depends on it. The routing bottleneck also explains the basin geometry that prior work observes without an underlying cause: safety is narrow along the gate direction because small perturbations there sever the connection between two otherwise stable representations. If fragility comes from depending on a few heads, then training the model to distribute that dependence should make it robust. We test this with noise-injected fine-tuning on the bottleneck heads. The model's dependence on them disappears, and the basin widens along that direction. Applying the same training to the top-attribution head changes nothing, because its contribution is redundant and removing it leaves the circuit intact.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.