Routing Can Hide What Unlearning Did Not Remove: A Forced-Routing Probe for Mixture-of-Experts
Abstract
In Mixture-of-Experts (MoE) language models, unlearning can lower forget-domain accuracy in two ways: by removing knowledge stored in expert parameters, or by changing which experts the router activates. Evaluation under the model's own routing conflates the two, scoring a removed capability and one the router merely hides alike. We introduce a forced-routing evaluation framework that separates them by also reading the forget domain through the experts that natural routing bypasses. With it, we show that natural-routing metrics can substantially overestimate unlearning, as capabilities that appear forgotten can remain recoverable through alternative expert paths. We further propose a recoverability-based expert selection criterion, which chooses experts by how much target capability remains recoverable through them rather than by how often or how strongly the router selects them. The two criteria select largely disjoint experts. Under representation-based unlearning, recoverability-based selection leaves less capability recoverable than SEUF's affinity-based selection at matched forgetting depth, with comparable retain utility. Unlearning in sparse models should therefore be evaluated by what the parameters retain, not only by what the default computation path exposes. Anonymous code is available at https://anonymous.4open.science/r/ForcedRoutingProbe.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.