Same Answer, Different Lesson: Defending On-Policy Distillation via Signal Decoupling
Abstract
On-policy distillation (OPD) poses a growing threat to proprietary large language models (LLMs), as a student can efficiently imitate the teacher under per-token supervision on its own trajectories. Yet existing anti-distillation defenses target teacher outputs, leaving OPD’s student-conditioned supervision underexplored. In this work, we address OPD protection in two components: First, we formulate OPD protection as a constrained adversarial optimization problem that decouples the teacher's prediction decision from the supervision signal exposed to the student, allowing the latter to be weakened while preserving the former. Second, we propose Selective Asymmetric Supervision Reshaping (SASR) to realize this objective at the OPD scoring interface. SASR consists of two key designs: 1) selective intervention, which focuses protection on positions with large teacher-student disagreement, 2) asymmetric reshaping, which modifies supervision according to the direction of the teacher-student gap. A decision-preserving projection further constrains the reshaped distribution to retain the teacher decision while controlling supervision distortion. Experiments across different teacher–student capacity settings and multiple OPD variants show that SASR consistently suppresses student performance, reducing accuracy by up to 7.26 percentage points compared with the evaluated state-of-the-art defense. To our knowledge, SASR is the first defense to explicitly target student-conditioned supervision in OPD, addressing a previously underexplored protection surface.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.