acceptodds
Under review as a conference paper at ICLR 2027

Where to Anchor? Choosing the Prompt Distribution for Alignment Anchors in Narrow Fine-Tuning

Abstract

Fine-tuning a language model on a narrow task can make it misbehave far outside that task, a problem called emergent misalignment (EM). Defenses keep the model close to its original behavior through its weights, training data, internal activations, or outputs; we ask which of these alignment anchors works best and how to compare them fairly. Such comparisons are easily distorted: a defense can lower EM just by learning the task less, and EM itself is scored by an automatic judge. We show that EM is a measurement-dependent quantity: even with the same frozen responses, changing the thresholds or selector can flip which defense appears best, and no alternative judge agrees closely enough with the primary judge to replace it. Human raters agree with the judge on 83% of responses, and most of their disagreements fall on responses scored right at its thresholds. We therefore propose a protocol that compares defenses only when they learn the task equally well, tunes each on questions kept apart from those used to test it, and reports how each result shifts with these choices. Under this protocol, penalizing how far the model's outputs drift from the original model's gives the lowest observed EM in all three model families tested. Computing this penalty on unrelated prompts lowers EM further, but only together with other changes; changing the prompts alone has no detectable effect. We also explain when unrelated prompts should help: if the same few internal factors drive misbehavior across tasks, an anchor on unrelated prompts can hold them in check without blocking the task.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.