acceptodds
Under review as a conference paper at ICLR 2027

MAxBench: A Multinomial concept recovery benchmark

Abstract

Concept localization and activation steering are increasingly common in interpretability research and its applications. Many existing methods implicitly assume binary or one-dimensional concepts, but many concepts are not binary: some contain many values, hierarchical structure, ordered values, among others. Recent work has explored multinomial concepts and found that their representation geometries are far more varied than those of binary concepts; this makes it unclear which methods or assumptions, and thus which localization and steering methods, are most appropriate. In this study, we introduce MAxBench, a geometry-agnostic evaluation framework for multinomial concept representations based on sampling from the recovered concept representation. Comparing localization methods across a range of geometry types, concepts, and models, we find the following: (i) where the concept representation is centered often matters as much as (and sometimes more than) the choice of basis and the dimensionality of the concept representation; (ii) affine subspaces outperform directions and linear subspaces (even at rank one); (iii) methods within the same geometry class tend to perform similarly; and (iv) no method consistently outperforms prompting, especially for larger models. These findings underscore the importance of expanding meta-evaluations of interpretability research to concepts with more varied internal structure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.