acceptodds
Under review as a conference paper at ICLR 2027

AlignRadar: Forecasting Emergent-Misalignment Risk Before Fine-Tuning

Abstract

Emergent misalignment (EM) is the unintended appearance of harmful, unsafe, or deceptive behavior after fine-tuning, including behavior outside the training dataset's domain. We study supervised fine-tuning (SFT), a widely used adaptation method in which task-specific datasets directly shape model behavior. Because cross-domain EM is usually discovered only after SFT and behavioral evaluation, a risky dataset may consume substantial compute and safety-review effort before the resulting model must be rejected or retrained. The central challenge is that inspecting a dataset before training does not directly reveal whether it will cause harmful behavior beyond its original domain. We introduce **AlignRadar**, a forecaster that operates before SFT using only a candidate dataset and an untouched base model. AlignRadar analyzes the dataset and the base model, uses patterns learned from completed fine-tuning runs, and returns a continuous score that ranks the dataset's EM risk. Tuned models and post-SFT evaluation results are used only to train AlignRadar and are never required when assessing a new dataset. Across 76 corpora from 13 dataset families, nested leave-one-family-out evaluation achieves **75.3% AUROC** and **76.1% average precision**. AlignRadar improves AUROC by **7.0 percentage points** over a base-model audit and by **21.6 points** over the same pre-SFT inputs trained without internal-change supervision. To evaluate cross-domain forecasting directly, we reconstruct outcomes using only prompts outside each corpus's training domain. AlignRadar achieves **72.2% AUROC** and **62.8% average precision**, improving over a direct EM classifier combined with the same audit by **14.2** and **9.2 percentage points**, respectively. These results show that learning from previous fine-tuning runs helps AlignRadar rank new datasets by EM risk before they are used for fine-tuning. Anonymized code is available at: https://anonymous.4open.science/r/AlignRadar_SFT_Anonymous-2BB1/

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.