acceptodds
Under review as a conference paper at ICLR 2027

To be or not to be: Mutually Exclusive Features in Language Models

Abstract

Sparse autoencoders have become a widespread tool for discovering interpretable features in language models and to intervene upon them. Yet the standard practice still treats features in isolation, leaving the structure of their relationships unexplored. We study these relationships through feature co-activation graphs, focusing on a particular suppressive relationship that we call mutual exclusivity: features that occur in similar contexts but not together. Across a wide range of models varying in size, we consistently find that mutually exclusive features form interpretable groups whose members represent alternative values of a common semantic type, such as numerals, pronouns, and months, to name a few. These groups exhibit consistent geometric structure: their decoder directions are highly similar, placing alternative values within narrow regions of activation space. Finally, we show that these features are causally relevant to model behavior: replacing a feature with another ME alternative predictably moves the model’s output toward the concept represented by the replacement feature. Together, our findings demonstrate that relationships among features reveal structure in model representations and mechanisms that are difficult to discover by studying them in isolation

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.