acceptodds
Under review as a conference paper at ICLR 2027

Lie Rates and Selective Censorship Dissociate During Distillation

Abstract

Can censorship pass from a teacher model to a student through training data screened to remove the censored topic? Prior work presents increased false answers about China as evidence of this "subliminal learning," even across model families. Yet that measure alone cannot distinguish selective censorship from a broader loss of accuracy. We test this distinction using a synthetically created benchmark comprised of matched questions about China and other jurisdictions, comparing students trained on filtered outputs from two censoring teachers with students whose training includes explicit China-related examples. Filtered training meaningfully raises instruction-tuned 3B students' false-answer rates, but none of seven filtered students show a detectable increase in censorship specific to China relative to their starting models. Adding explicit China-related examples consistently increases this selective censorship measure, while false-answer rates do not consistently rise. At 120B parameters, we measure selective whitewashing: minimizing or omitting sensitive facts disproportionately for China, which we find explicit exposure increases as well relative to filtered training. Higher false-answer rates alone therefore do not establish that censorship has transferred subliminally across model families.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.