acceptodds
Under review as a conference paper at ICLR 2027

Why Benign Post-Training Induces Safety Gate Collapse: Changes in Harmful Activation Geometry Relative to the Inherited Refusal Gate

Abstract

Benign post-training can drive an aligned model from high refusal and low attack success to the opposite pattern, we call this transition Safety Gate Collapse. We introduce Class-Conditional Activation Geometry to quantify differences between harmful and benign activation changes, and the Inherited One-Sided Refusal Gate to connect these geometric changes to the model's refusal behavior using a fixed boundary calibrated on the aligned model. We find that (i) benign SFT on code or general instructions and self-play RLVR lower refusal on eight of the nine post-trained models we study, by up to %, and raise attack success on the same eight; (ii) harmful and benign requests remain linearly distinguishable from internal activations (probe AUROC % on every post-trained model), while occupancy of the inherited refusal region by harmful activations on Llama-3.1-8B-Instruct falls from to after code SFT and to after Alpaca SFT; (iii) the sign of the harmful centroid's displacement along the gate normal matches the direction of refusal change on all nine models, whereas total displacement does not; and (iv) steering along the inherited gate direction raises refusal and lowers attack success on every post-trained model, but can also increase over-refusal on benign requests. These findings identify changes in harmful activation positions relative to the inherited gate as a diagnostic of refusal degradation, even when harmful and benign requests remain linearly distinguishable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.