acceptodds
Under review as a conference paper at ICLR 2027

Concept-Conditioned Scores Can Rank Sparse Dictionaries Without Using the Concepts

Abstract

Sparse dictionaries are increasingly compared by scores computed from concept labels: how compactly a labeled concept is represented, how well it matches a feature, how much information a feature carries about it. Concept labels give such a score an apparent semantic grounding. Whether the ranking of dictionaries it induces depends on the labels is a separate question, and existing sanity checks, which randomize the dictionary, do not ask it. On dictionaries we train for chest radiographs and on public sparse autoencoders for language models, under three label systems, replacing each concept's support with a size-matched random set of documents leaves the ranking under most scores unchanged. The invariance reaches model selection: on a controlled sweep, compactness scores rank collapsed dictionaries first and give their best value to a dictionary trained on permuted concept targets, whereas probes favor real targets on every one of 20 paired seeds. We therefore propose TWIN (two-way identity null), a protocol for auditing a concept-conditioned score before it is used to compare dictionaries: it destroys concept identity once at evaluation and once at training, controls occupancy, and anchors every score to probing. Applied to fourteen scores, TWIN separates failures of different kinds: rankings that do not move when the concepts are randomized, preferences for real targets that vanish once the feature budget is matched, and a matching score whose ranking follows the live count. The mutual-information gap is the most label-sensitive score, and its sensitivity to discretization keeps it a diagnostic. For compactness the failure can be explained: the score depends only on the sorted shape of a concept's activation profile, that shape stays close to the label-free marginal one on all our pools, and for the primary score we prove that this closeness confines its value to an interval around the value it takes with no labels at all.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.