PACO-SAE: Evaluating Concept Grounding and Feature Splitting in Vision Sparse Autoencoders
Abstract
Sparse autoencoders (SAEs) are used to interpret vision models, but predicting a concept does not establish whether a feature activates on the object it represents. We introduce PACO-SAE, an object-level benchmark and evaluation protocol that evaluates concept grounding as semantic segmentation. Built from exhaustively annotated PACO-LVIS objects, it combines grounding F1 with separate feature selection and testing across 63 object concepts. Across four SAE families and four sparsity levels, image-level probing and peak localization select features with lower test grounding F1 than grounding-based selection, showing that neither substitutes for evaluating the entire activated region. PACO-SAE also enables object-level analysis of feature splitting: different features best ground different objects of one concept, substantially outperforming a single feature selected for the highest average grounding. To examine whether these splits reflect visual specialization or residual splitting, we introduce Pairwise Visual Separation (PVS), which compares visual differences between groups of objects with variation within each group, without additional subclass labels. A human study of 200 feature pairs supports PVS as an indicator of recognizable visual differences. Together, PACO-SAE establishes an object-level evaluation foundation, reveals the limits of prediction and localization as grounding criteria, and connects the extent of feature splitting to its visual meaning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.