Targeted Routing-Aware Calibration for Machine Unlearning in Mixture-of-Experts Language Models
Abstract
Machine unlearning aims to remove the undesired data from large language models (LLMs) while preserving model utility. As Mixture-of-Experts (MoE) LLMs become prevalent, there is a growing need for effective unlearning methods that account for their distinctive architecture. In this work, we observe that a small subset of experts has substantially higher activation frequencies on forget data than the other experts, while generic retain data activates these experts much less frequently. These observations motivate considering both which experts to update and how to direct retain regularization toward them. We propose , argeted outing-ware alibration for xperts, a two-stage framework for MoE unlearning. TRACE first selects experts with high activation frequencies on a small calibration subset of forget data. During unlearning, it reweights the retain loss to prioritize retain samples whose activation frequencies on the selected experts more closely match those of the forget data, providing targeted retain regularization. Experiments across multiple MoE LLMs show that TRACE preserves substantially higher utility at comparable forgetting levels, including a 4.55-point MMLU improvement over the strongest external baseline on Qwen1.5-MoE-A2.7B-Chat and a 3.96-point improvement on DeepSeek-V2-Lite-Chat.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.