Tracing Subliminal Learning Through Pretraining
Abstract
Large language models trained on data generated by other models can sometimes inherit traits unrelated to that data, a phenomenon called subliminal learning. Because future models will increasingly be trained on model-generated text, it is important to understand the conditions under which this transfer occurs. Using Pythia, PolyPythia, and DataDecide OLMo models, we isolate how weight initialization, pretraining data order, pretraining data corpus, model size, and architecture affect subliminal transfer. We study both hard label distillation, the standard case where the student model is trained on numbers generated by a biased teacher, and on-policy distillation, where the student generates the data and the teacher's logits provide the learning signal. Transfer persists when the teacher and student base models differ in any and all of these factors under both settings. It is also strongly asymmetric: swapping which base model acts as teacher and which as student can substantially change the magnitude of transfer. Finally, we find that across pretraining checkpoints, the ability to transmit a trait and the ability to acquire one develop separately, so some checkpoints are effective teachers but weak students. Hard-label transfer is detectable but requires more extreme conditions, while soft-label transfer is much larger, making it the more realistic concern.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.