Token-Selective On-Policy Self-Distillation for LLM Unlearning
Abstract
Large language model (LLM) unlearning aims to suppress designated knowledge while preserving non-target capabilities, without retraining from scratch. However, existing parameter-updating methods often rely on externally constructed supervision, while directly distilling in-context unlearning (ICU) may also transfer prompt-induced changes unrelated to forgetting. In this work, we propose Token-Selective On-Policy Logit Distillation (TOLD), a response-label-free unlearning method that uses ICU as supervision for parameter updating. TOLD compares a frozen teacher's predictions with and without the ICU prompt along student-generated trajectories. It then selectively transfers token-level changes for suppression and promotion, limiting unrelated changes during distillation and enabling prompt-free inference. Experiments on the RWKU benchmark across two backbones show that TOLD reduces mean forget-set recall by 39.6-42.5% relative to the base models while largely preserving neighboring knowledge and general utility. TOLD also demonstrates robust forgetting across nine adversarial attacks. Additional experiments on MUSE-Books extend these findings beyond entity-level unlearning. These results demonstrate the effectiveness of selective on-policy distillation for prompt-free behavioral unlearning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.