WHY SUBLIMINAL LEARNING NEEDS SO MUCH DATA: A NOISY INVERSE VIEW THROUGH STEERING VECTOR RECOVERY
Abstract
Subliminal learning allows a student model to inherit a teacher’s behavioral trait from semantically unrelated generated data, but existing demonstrations often require tens of thousands of independent carrier examples. We ask whether this scale is needed because small datasets fail to cover the trait signal, or because sampled tokens are hard labels and provide a noisy representation of that signal. Using subliminal steering as a controlled testbed, we represent the teacher trait by a known vector \(\Delta_T\), recover a student vector \(\Delta_S\), and calculate their cosine similarity to measure subliminal transfer effect. To isolate the effect of hard-label noise, we compare token-level NLL supervision with full-distribution KL supervision on matched carrier prefixes. The results create a two-sided contradiction for the two candidate explanations. With sufficiently many carriers, the hard- and soft-supervised mean gradients are nearly collinear and have the same low alignment with . From this gradient analysis, hard-label noise appears to be readily averaged out and therefore should not affect the final subliminal transfer. Under iterative optimization, however, soft supervision recovers almost perfectly from only hundreds of examples, whereas hard supervision remains substantially limited even with tens of thousands. The large gap between their endpoints shows that hard-label noise is not average out, and nevertheless decisive for final recovery. At the same time, the success of small-data soft supervision rules out insufficient carrier coverage as the primary reason that hard-label learning requires large datasets. Thus noise appears negligible in the local gradient but consequential over optimization, while the carrier already contains enough information for recovery at a much smaller scale. We explain this discrepancy by formulating subliminal learning as a noisy inverse problem. Locally, the carrier task Fisher matrix transforms the teacher trait direction, and the carrier learning gradient for becomes: \(F(\Delta_T-\Delta_S)\), causing hard and soft gradients to share the same distorted direction, which explains the poor alignment with regardless of carrier data size. Iterative optimization progressively inverts this transformation, allowing soft supervision to recover the teacher, but simultaneously amplifying the small residual error in hard supervision. Thus large carrier datasets are not primarily required to fully catch the trait signal; they reduce discretization error amplified by Fisher inversion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.