acceptodds
Under review as a conference paper at ICLR 2027

MixCoder: Decomposing Neural Computation into Low-Rank Transformations over Low-Dimensional Representations

Abstract

Fine-grained circuit analysis seeks to decompose neural computation into interpretable units and describe their interactions with attribution graphs. Existing feature-based approaches typically model semantic representations as isolated directions and computation as interactions among features, binding what is represented and how it is transformed within the same feature-level description. Growing evidence instead suggests that neural representations can occupy structured low-dimensional geometries, while effective computation over them can take the form of low-rank transformations. Motivated by this view, we argue that mechanistic interpretability should explicitly separate activation patterns, representations, and transformations, and decompose model computation into **low-rank transformations of low-dimensional representations**. From a generative model that factorizes these roles, we derive **MixCoder**, which identifies low-dimensional representation modules, lets representation states determine sparse activation patterns, and uses explicit operators to model transformations over them. Applied to language models, MixCoder reveals which representations are jointly read, how they are transformed, and where the results are written back. We further use this framework to study **transformation interference** when multiple low-dimensional transformations share limited computational space and its effects on model performance, providing insights into improving performance through mechanistic interpretability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.