Misaligned model variants serve as an effective source of training for harm detection probes
Abstract
The ability to detect misaligned outputs from large language models (LLMs) is critical for their safe deployment. Linear activation probes have emerged as an effective solution, particularly when incorporated in a multilayered monitoring pipeline (Cunningham et al., 2026). They are fast, cheap, and easy to train and deploy. However, these probes can also be brittle, particularly when trained on off-policy data—data generated by a source other than the model itself (Kirch et al., 2026). On-policy data, however, are hard to source because typical model generations are mostly benign. Here, we propose a pipeline where probes are trained on outputs from misaligned model variants, derived from the base model via finetuning or other interventions. We use harm generation as a sample do- main. We find that probes trained on activations from harmful model variants—refusal-ablated, emergently misaligned, and broadly harm finetuned—can effectively identify harmful generations by the base model, outperforming off-policy probes. We also systematically vary different design choices during probe training (outputs selected as positive vs. negative labels and the model supplying activations for probe training), showing that these choices can strongly affect the representation learned by the probe and its ability to classify harmful out- puts. In particular, we show that the typical probe training setup, where benign responses to benign prompts are contrasted with harmful responses to harmful prompts, is brittle compared to contrasting benign versus harmful responses to harmful prompts. Overall, we provide evidence that misaligned model variants are an effective source of data for probe training and outline a comprehensive set of guidelines for better probe design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.