SafeOPD: Safety-Preserving On-Policy Distillation from Domain-Specialized Teachers
Abstract
On-policy distillation (OPD) has emerged as an effective paradigm for transferring advanced reasoning capabilities from stronger teachers to smaller students. We find, however, that even a small fraction of harmful training prompts can substantially weaken the safety alignment of an already aligned student, despite continued gains in mathematical and medical reasoning capabilities. We term this failure mode on-policy safety forgetting. To address this problem, we introduce SafeOPD, a general OPD framework for safety-preserving post-training of large language models (LLMs). SafeOPD incorporates two key designs. First, we identify safety-relevant states using a small calibration set, locating where teacher supervision may conflict with the student's existing safety behavior. Second, we selectively constrain teacher supervision at these states using the initial aligned student, while preserving the original distillation signal elsewhere. This local intervention preserves the student's safety behavior without broadly weakening capability transfer. Experiments across the domains of mathematical and medical reasoning show that SafeOPD preserves the initial student’s safety alignment while retaining most of the capability gains obtained from the corresponding specialized teacher. Overall, our results highlight the importance of incorporating safety preservation directly into the distillation process, rather than treating safety alignment as a separate post-training stage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.