LEAF-OPD: Learning to Forget from Privileged Contexts via On-Policy Distillation
Abstract
LLM unlearning aims to erase targeted knowledge from large language models while preserving their utility. However, many existing methods focus on suppressing target knowledge without explicitly teaching the model how to respond to queries about it. We propose LEAF-OPD (LEArning to Forget from Privileged Contexts via On-Policy Distillation), an LLM unlearning algorithm in which a frozen teacher policy provides token-level supervision on responses generated by the student policy. On forget data, the teacher policy uses privileged contexts (additional instructions not to reveal knowledge about forget targets) to guide forgetting. On retain data, the same frozen teacher policy provides supervision without privileged context, encouraging the student to preserve the teacher’s responses and thus maintain model utility. The student policy operates without privileged context in both cases. Instead of simply pushing the model away from the forget data distribution, LEAF-OPD explicitly teaches the student how to behave when encountering knowledge that should be forgotten. A key challenge is that the teacher usually provides effective supervision only at the first few tokens along the student's trajectories, with little effective supervision at subsequent tokens. To characterize the teacher's supervision strength, we introduce Supervision Signal (SS), which measures the strength of corrective supervision over a specified token interval. We evaluate SS over the entire trajectory (SS₁) and its second half (SS₀.₅) to assess the overall strength and late-stage persistence of corrective supervision, respectively. When supervision is weak, LEAF-OPD applies lightweight SFT to strengthen the teacher's context-conditioned refusal behavior. This can substantially increase both SS₁ and SS₀.₅, transforming a bad teacher into a good teacher. We evaluate LEAF-OPD on three widely used LLM unlearning benchmarks: TOFU, RWKU, and MUSE-News. LEAF-OPD achieves the best performance regarding Forget Quality, Model Utility, and Forget ROUGE-L on TOFU, while maintaining competitive retain performance. On RWKU, LEAF-OPD outperforms the second-best method by 34% in Forget Quality while maintaining competitive Retain Quality. These results demonstrate that privileged-context supervision with on-policy token-level distillation provides an effective approach to LLM unlearning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.