acceptodds
Under review as a conference paper at ICLR 2027

DEX: Data-Emergent Experts for Capability Ablation in Large Language Models

Abstract

Modern language models are trained on vast corpora that inevitably contain harmful or dual-use content, and a capability that helps a legitimate user can just as easily assist a malicious one. A safe deployment therefore benefits from the ability to strip a specific capability from a trained model on demand, without paying to retrain and without degrading everything else the model does. We propose DEX (Data-Emergent Experts), a dense model inspired by the Mixture-of-Experts architecture but with an embedding-based router instead of a learned one. We first cluster the training corpus by each document’s meaning using a pretrained embed- ding model, letting the data’s own structure decide both the categories and which documents belong to each, rather than relying on a predetermined set of topics and fixed assignments. Each cluster is assigned its own expert, and unlike standard sparse routing, all experts stay enabled: every non-ablated expert contributes to each input, weighted by a softmax over the similarity between the input embed- ding and the cluster centroids. Removing a capability is then as simple as zeroing the corresponding experts. We study how to balance capability removal against the retention of knowledge we want to keep, and we find each cluster captures a specific interpretable specialization. DEX forgets far more of the target capability while marginally affecting the rest. Additionally, since knowledge is split into finer clusters, DEX can remove narrower subspaces where prior methods can only drop whole domains.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.