acceptodds
Under review as a conference paper at ICLR 2027

Router-Free Modular Storage for Knowledge Unlearning in Large Language Models

Abstract

Machine unlearning aims to remove the influence of selected training data sources, such as individual users or domains, without retraining a model from scratch. This is difficult because information from many sources is entangled in shared model parameters, so that removing one source harms utility on others. Unlearning by construction sets up training to avoid entanglement by confining information on individual data sources to separate modules of a trained model. However, this causes deployment overhead because it requires a router and source labels for every data point at inference time to determine which module is responsible. To overcome this limitation, we propose Modular Unlearning via Self-Routing (MUSR), which places one small module per source in parallel with the MLP layers and passes every input through all of them. During training, each module learns to contribute on its own source and to stay near zero on every other. Thereby, the modules gate themselves and inference is a sum of their outputs with no router and source labels. Our empirical evaluation across four different language models and across pretraining and fine-tuning setups highlights that MUSR matches the removal performance of the modular baselines while retaining 99.7% of its pre-unlearning model utility on average. Beyond single deletions, MUSR supports repeated deletion requests and the addition of new sources after training. Overall, MUSR makes unlearning by construction more practical for standard model inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.