ModULE: Learning Modular Expert Representations for Localized Machine Unlearning
Abstract
Machine unlearning aims to remove the influence of specified training data while preserving performance on retained data. Existing methods typically modify models post hoc and often require broad parameter updates, resulting in substantial computational cost and collateral degradation. We study whether structured deletion requests can instead be made localizable to small expert subsets. We introduce Modular Unlearning via Localized Mixture of Experts (ModULE), a deletion-aware MoE framework that promotes routing-space localizability. During learning, ModULE combines sparse routing, balanced expert utilization, and centered kernel alignment to encourage sparse and differentiated expert pathways. Given a deletion request, it selects experts that provide high forget-route coverage with limited retain exposure and identifies retained examples whose routes intersect the selected experts. Unlearning then updates only the selected experts, while output and route distillation protect exposed retained examples and limit layer-wise routing shifts. Under hard routing and frozen shared parameters, retained inputs whose pathways avoid the updated experts remain unchanged. Experiments on structured class- and domain-level deletion, with random deletion as a stress test, show that ModULE improves the forgetting-utility trade-off while updating fewer parameters than matched dense and conventionally trained MoE baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.