Identifying the Effect of Training Context in Emergent Misalignment
Abstract
Supervised fine-tuning adapts language models to new tasks, but narrow training can also induce misaligned behavior beyond the training domain. Understanding this phenomenon requires isolating the effect of training context and examining how it is expressed by the resulting parameter change. Building on target-matched controls, we formulate target-matched SFT, holding assistant targets fixed while identifying the training-context effect within paired training blocks under a common evaluation protocol. The base model provides a reference for separating the contrast between the fine-tuned models from the change under each condition. Across four paired training blocks of Qwen and Mistral, insecure training increases coherent low-alignment tail risk by 2.25–6.25 percentage points relative to educational training. We find interpolation along the learned parameter contrast yields positive tail-risk trends, while spectrum-matched orientation controls produce weaker sequence-distribution and behavioral responses. The behavioral effects also vary across evaluation panels and base models: unrelated-code effects have opposite signs in two models. These results identify a training-context effect. Parameter orientation shapes its functional response, and its measured expression varies with the base model, outcome, and evaluation panel.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.