acceptodds
Under review as a conference paper at ICLR 2027

Canary-Controlled Feedback Interleaving for Emergent Misalignment in LLM Fine-Tuning

Abstract

Fine-tuning large language models (LLMs) on narrowly defined tasks can lead to emergent misalignment, where undesirable behaviors appear on unrelated inputs even when the training data itself is benign. A common mitigation strategy is to interleave auxiliary safe data during training; however, existing approaches typically adopt fixed mixing schedules, which are unable to respond to the evolving nature of misalignment. We introduce a canary-controlled feedback interleaving approach that adjusts data mixing during training based on online risk signals. Specifically, a set of canary prompts is used to probe model behavior, producing a running estimate of misalignment risk via exponential smoothing. This signal drives a simple decision policy with hysteresis, allowing the training process to increase or relax intervention in a stable manner as conditions change. Evaluations on the Security EM benchmark with Qwen2.5-7B-Instruct show that the proposed approach reduces the misalignment rate from 7.15% to 5.39% (a 25% relative decrease) compared to fixed-ratio interleaving, while also achieving 4× lower variance across random seeds. Additional studies reveal that performance gains are primarily attributable to adaptive intervention timing, rather than the overall proportion of safe data, as fixed-schedule variants perform 36% worse under the same data budget. These findings indicate that effective alignment during fine-tuning depends less on static data composition and more on responsive adjustment during training, suggesting a shift toward adaptive strategies for improving robustness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.