acceptodds
Under review as a conference paper at ICLR 2027

Attributing Class Recognition to Invariants with Recursive Peeling

Abstract

A sparse autoencoder splits what a vision model computes into concepts, and most work studies these concepts one at a time. We ask instead what they reveal together. We call them class unique when they fire consistently on one class and shared when they fire consistently on several, representing invariants of different sizes. We propose Multi-level Clustering with Peeling (MCP) to uncover how these invariants organize classes and contribute to recognition. MCP clusters classes at resolutions determined by how widely their concepts are shared. Concepts shared only within one cluster are peeled as cluster unique concepts, exposing further relationships at subsequent levels. Applied to the final-layer class token of the self-supervised DINOv2 on ImageNet, MCP reveals a forest of visual categories that departs from the WordNet hierarchy. The clusters are tighter and more semantically coherent than size-matched random clusters. Ablations with a frozen linear classifier show that removing the concepts peeled from a cluster lowers accuracy within that cluster, with near-zero median effects outside it. Different classes rely on class unique invariants, cluster unique invariants, or both. Removing both kinds produces super-additive accuracy losses, while either kind alone can preserve recognition. These results suggest that the model recognizes classes by parallel abduction, weighing evidence for single classes and for clusters of related classes at once. Class unique invariants distinguish individual classes, while cluster unique invariants support several competing classes. Unlike human vision, which reaches its decision by recursive abduction, the model weighs every candidate in a single pass.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.