Where to Anchor? Choosing the Prompt Distribution for Alignment Anchors in Narrow Fine-Tuning
Abstract
Fine-tuning a language model on a narrow task can make it misbehave far outside that task, a problem called emergent misalignment (EM). Defenses keep the model close to its original behavior through its weights, training data, internal activations, or outputs; we ask which of these alignment anchors works best and how to compare them fairly. Such comparisons are easily distorted: a defense can lower EM just by learning the task less, and EM itself is scored by an automatic judge. We show that EM is a measurement-dependent quantity: even with the same frozen responses, changing the thresholds or selector can flip which defense appears best, and no alternative judge agrees closely enough with the primary judge to replace it. Human raters agree with the judge on 83% of responses, and most of their disagreements fall on responses scored right at its thresholds. We therefore propose a protocol that compares defenses only when they learn the task equally well, tunes each on questions kept apart from those used to test it, and reports how each result shifts with these choices. Under this protocol, penalizing how far the model's outputs drift from the original model's gives the lowest observed EM in all three model families tested. Computing this penalty on unrelated prompts lowers EM further, but only together with other changes; changing the prompts alone has no detectable effect. We also explain when unrelated prompts should help: if the same few internal factors drive misbehavior across tasks, an anchor on unrelated prompts can hold them in check without blocking the task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.