Coordinate Concentration of Safety Alignment in Low-Rank Adapters
Abstract
Fine-tuning a safety-aligned language model on benign data is enough to degrade its refusal behavior. The established defense projects the low-rank update away from an alignment-sensitive subspace, which assumes that the update has to be steered around safety rather than kept out of it. We ask where that subspace sits among the coordinates of the adapter itself. If safety energy concentrates on a small fraction of coordinates rather than spreading across all of them, those coordinates can be identified once and held fixed during any subsequent fine-tuning, replacing the projection with a simpler intervention that does not depend on the downstream task. To answer this, we approximate the curvature of the loss on harmful prompts and measure how sensitive the model's safety behavior is to changes at each adapter coordinate: coordinates where the loss curves sharply are the ones whose movement during fine-tuning would most disrupt alignment. Across three model families, we find that the distribution of safety-sensitive energy is heavy-tailed, and one tenth of the coordinates carries to of it, a property we call coordinate concentration. Holding that tenth fixed during fine-tuning preserves refusal behavior across four base models and two harmful-prompt benchmarks, performing comparably to three published defenses while requiring nothing of the downstream user beyond attaching the mask. The concentration also explains why targeted selection matters even at a heavy lock rate. A random mask of equal size locks nearly the same coordinates as a targeted one, yet performs substantially worse, because the small set of coordinates each leaves free to train are almost entirely different, and safety is determined by that free set.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.