acceptodds
Under review as a conference paper at ICLR 2027

Anchored but Not Protected: When Refusal Erodes During Fine-Tuning, and Whether Holding One Direction Stops It

Abstract

Fine-tuning an aligned chat model on ordinary instruction data makes it more willing to answer harmful requests. One family of defences picks a direction in activation space, calls it the refusal direction, and penalises movement along it while training runs. We audit that idea on LLaMA-3-8B-Instruct and Qwen2-7B-Instruct. Erosion is concentrated: 52 of the 517 optimiser steps in epoch 1 carry 89 to 101 percent of the net displacement that epoch produces, under four learning-rate schedules. The coordinate these defences hold mediates the refusal that fine-tuning removes. On four exploratory harm sets, restoring its vanilla value at inference takes undue compliance down 18 points to within 2.2 of the vanilla model, and past vanilla under a reference classifier, while the same operation on a random direction does nothing measurable; by ten epochs it recovers about three quarters of the gap. What fails is the calibration. On LLaMA-3, at the strength we inherited, the penalty prevents 9.61 degrees of drift where our acceptance bar needs 15.18 degrees. Anchoring on a direction pooled over nine refusal categories prevents 3.08 degrees more on LLaMA-3 and 2.91 degrees more on Qwen2, lowers undue compliance on LLaMA-3 by 2.0 to 3.6 points across three readings of one rubric, and still misses the bar. Tripling the penalty on LLaMA-3 clears it and cuts undue compliance on the harm sets by a further 6.75 and 8.25 points under the primary scorer and the reference, at a cost of one extra false refusal in 250 and no loss of ReClor accuracy relative to training with no penalty at all. Underneath sits a measurement problem large enough to decide results: three scorers disagree by up to 70 points on identical responses, and one published rubric built two defensible ways disagrees with itself on 6 percent of them, moving the pooled anchor's pre-registered effect from 3.6 points to 2.2 and changing whether it resolves.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.