HPOD: Hierarchical Preference Optimization via On-Policy Distillation for LLM Concept Unlearning
Abstract
Large Language Models (LLMs) achieve remarkable capabilities from massive corpora but may also memorize harmful or undesired knowledge. *LLM unlearning* removes the influence of such knowledge without retraining. Existing methods mainly focus on instance-level forgetting, yet sample removal may not eliminate related knowledge across semantically connected instances. This motivates *LLM concept unlearning*, which suppresses target concepts while preserving non-target knowledge and general capabilities. However, it faces two key challenges: behavior-aligned supervision under evolving generation distributions and coordination of conflicting forgetting and retention objectives. To address these challenges, we propose **HPOD**, a hierarchical preference optimization framework via on-policy distillation. **HPOD** first leverages student-generated trajectories to perform on-policy distillation, providing trajectory-level guidance to preserve non-target knowledge while suppressing target concepts through reverse optimization. Then, a reward-guided reinforcement optimization strategy promotes forgetting-oriented behaviors by calibrating model responses while maintaining coherent generation. Finally, an adaptive multi-objective coordination strategy dynamically balances competing objectives and more effectively mitigates optimization conflicts through gradient coordination. Extensive experiments on real-world datasets demonstrate that **HPOD** achieves effective forgetting while preserving model utility. The code is available at https://anonymous.4open.science/r/HPOD-0004/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.