Regularization Amplifies Subliminal Learning
Abstract
Today’s language models are trained under the influence of other language models—through data generation or data filtering—to transfer preferences, knowledge, and skills. However, recent work has shown that hidden biases can also be transferred through data with no apparent semantic relationship to those biases. Can regularization limit this covert influence? By limiting adaptation to the training data, regularization might be expected to suppress behavioral changes on seemingly unrelated inputs. We systematically study this hypothesis across subliminal learning, phantom transfer, and logit-linear selection. We find the opposite: *regularization amplifies the transfer of hidden biases*. The effect is particularly dramatic in subliminal learning: at small data scales, low-rank LoRA produces strong transfer while high-rank LoRA and full fine-tuning produce none. Two classical means of mitigating overfitting let the larger students acquire the bias as well. More unique training data lets both high-rank LoRA and full fine-tuning acquire the bias, and L1 or L2 penalties do the same for high-rank LoRA. Phantom transfer and logit-linear selection show similar trends. Together, these findings suggest that hidden bias transfer benefits from the simplicity biases that regularization induces. Our results reveal that seemingly benign training choices can play an important role in amplifying the transfer of hidden biases between models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.