acceptodds
Under review as a conference paper at ICLR 2027

RExU: Recovery-Expert Unlearning for Mixture-of-Experts

Abstract

Machine unlearning aims to remove target knowledge while preserving other capabilities. Existing evaluations measure forget efficacy (FE) under normal inference, treating low FE as evidence of successful forgetting. This assumption can fail for Mixture-of-Experts (MoE) models, where each input activates only a sparse subset of experts: an unlearning method may suppress target knowledge along the natural route - the model's default expert selection without inference-time intervention - while leaving the same knowledge recoverable through alternative experts. We expose this off-route recovery through counterfactual routing and introduce Worst-Route Forget Efficacy (WRFE), which measures the maximum forget-set recovery over an audited route family including the natural route. Statistically above-chance recovery under such rerouting provides direct evidence that supposedly forgotten knowledge remains accessible in the model. Empirically, we find substantial off-route recovery despite successful natural-route forgetting. On WMDP-Cyber, for example, GA+SEUF lowers FE by percentage points, from to , under natural routing, yet bypassing the edited expert raises FE by percentage points to . Thus, natural-route FE alone can substantially underestimate residual recoverability in MoE models. To mitigate this failure, we propose Recovery-Expert Unlearning (RExU), which iteratively identifies recovery experts contributing most to worst-route recovery and selectively unlearns them. Across three MoE models, four base unlearners, and two benchmarks, RExU substantially reduces WRFE while maintaining competitive natural-route FE and retain utility. Our results motivate jointly evaluating MoE unlearning through natural-route FE, WRFE, and retain utility, and establish RExU as a practical method for improving this trade-off.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.