DUNE: Disentangling Polysemantic Neurons Without Losing Predictive Performance
Abstract
Deep neural networks (DNNs) are widely used, but understanding what they actually learn remains difficult. A major obstacle is that individual neurons often respond to multiple unrelated concepts, obscuring the network's internal decision process. Existing approaches, such as sparse autoencoders, can disentangle these mixed signals into more meaningful, “monosemantic” features, but typically introduce approximation errors that may degrade downstream performance. Consequently, these methods exhibit a Pareto frontier between monosemanticity and functional faithfulness. To address this, we introduce _DUNE_ (**d**isentangling **u**nits with **n**o **e**rror), a method that decomposes DNN neurons into more monosemantic subunits while preserving their original function by design. DUNE identifies concept-specific contributions from the preceding layer and re-routes them into separate subunits whose aggregation exactly recovers the original activation. DUNE can be applied directly to pretrained models, without labels or training of an auxiliary model. Across several models, including DINOv2, DUNE achieves state-of-the-art or competitive monosemanticity while exactly preserving functional faithfulness, substantially advancing the existing Pareto frontier between these properties. DUNE operates efficiently and enables fine-grained interventions such as concept-specific steering.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.