How to Remove Subliminal Signals?
Abstract
Subliminal signals are subtle statistical patterns in model-generated data that can transmit information about the teacher even when individual examples appear benign. We use the *teacher state* to denote any source-specific factor that affects the teacher's generation distribution, such as a behavioral trait, system prompt, or training or inference-time intervention. We study how to remove information about this state when the state itself is unknown, no clean reference model is available, and the downstream training procedure is not under our control. Our approach uses verified task-preserving transformations that change the representation of an output without changing its task-relevant content. We introduce *Self-Sanitized Generation (SSG)*, a data-side sanitization procedure that uses only samples and response probabilities from the original teacher. SSG removes state information at the distribution level, including signals that survive ordinary rewriting. Under a standard exponential-tilting model in which the teacher state reweights the probabilities of generated outputs, the sanitized distribution becomes exactly independent of the state. We further characterize when exact cancellation is possible beyond this model and quantify the remaining state dependence when exact removal cannot be achieved. Finally, we validate SSG on a storytelling task in which the subliminal signal is a preference for a specific animal, injected via system instructions on Qwen3.5-9B, with paraphrasing as our task-preserving transformations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.