Defence against the dark arts of distillation: Unsupervised discovery of subliminal learning
Abstract
When a language model is distilled from a teacher that shares its base parameters, the student can inherit behavioural traits of the teacher—including misalignment—even from data that is semantically unrelated to the trait (Cloud et al., 2026). All existing defences require knowing the trait. We remove this requirement and prove that a subliminally transferred trait can be recovered from the parameter updates alone, without supervision, by exploiting a single asymmetry: the trait is expressed consistently across diverse datasets generated by the teacher, whereas task-specific and shared capabilities vary. The resulting estimator returns a weight-space direction that can be amplified to discover and interpret the trait or projected out to suppress it. Applied to the subliminal-learning benchmarks of Cloud et al., 2026, our method recovers the transmitted animal-preference traits without any trait labels and suppresses their expression with almost no impact to task performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.