acceptodds
Under review as a conference paper at ICLR 2027

Sparse Compositionality Identifies Features of Computation

Abstract

Does a reconstruction objective recover the features that participate in a network's computation, or a basis that merely reconstructs its activations? The distinction matters if, as a growing body of work suggests, learners generalize through compositional sparsity. The features that explain such a network are then the intermediate functions it composes, and nothing guarantees that a featurizer trained to reconstruct will single them out. In this work, we make the distinction measurable by planting a ground-truth feature DAG, whose intermediate features and compositional relations are known, inside a transformer through per-layer supervision. Reconstruction-based featurizers, including sparse autoencoders, do not reliably identify these intermediate features as individual latents, although the same architectures recover them with direct supervision. This suggests that reconstruction alone cannot identify the planted basis even when the featurizer has the capacity to represent it, isolating a critical identifiability failure. We address this by constraining the featurizer dictionary to represent features relevant to downstream computation by pressuring the features to align with the Jacobian. On the synthetic model, this featurizer identifies the ground-truth basis to reveal the planted compositional hierarchy where baselines could not. We then apply this approach to DINOv3, where no intermediate ground-truth is available. Without any supervision, the resulting attribution graphs yield part-to-whole object circuits: features for a rabbit’s eye, nose, snout, and chest funnel into a downstream rabbit feature. Intervening on targeted upstream features changes the responses of connected downstream features, supporting the functional relevance of these links. For instance, patching a stripe concept on the horse body shifts a frozen ImageNet readout toward the zebra. Together, our results show that grounding features in downstream gradients addresses shortcomings of optimizing reconstruction alone, providing a path from featurizing representations to featurizing computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.