acceptodds
Under review as a conference paper at ICLR 2027

Back to the basics: Interpreting Vision Embeddings via Binary Sparse Autoencoders

Abstract

Sparse autoencoders (SAEs) have emerged as the primary framework for interpreting the internal representations of vision (and language) models, decomposing dense, entangled embeddings into sparse, more human-interpretable latent codes. Continuous activations trained under sparsity constraints have come to dominate this line of work, due to their manageable optimization dynamics. In contrast, latent codes remain scarce within the SAE setting, despite offering a natural and arguably unambiguous notion of concept presence. Binary SAEs trained with Gumbel relaxations or straight-through estimators are prone to training instability and typically underperform their continuous counterparts. In this work, we argue that binary SAEs are not intrinsically flawed: rather, it is they are trained that directly and disproportionately impacts the quality of their representations. To address this, we introduce (inary tom election by terative oordinate wapping), a binary SAE that draws inspiration from classical dictionary learning by casting binary code inference as a Gauss–Seidel local search over supports, in which every accepted coordinate swap strictly decreases the reconstruction objective. We compare BASICS-SAE against Gumbel-relaxed and Local Winner-Take-All binary SAEs, and against continuous TopK and BatchTopK SAEs, on CUB and ImageNet using CLIP and DINOv3 embeddings. Beyond reconstruction fidelity, we evaluate dictionary quality using downstream few-shot classification and formal mechanistic interpretability metrics, based on recently proposed concept alignment scores. BASICS-SAE is competitive against continuous baselines in reconstruction while substantially outperforming all binary and continuous models in downstream code quality and concept alignment in most settings. Our results establish binary SAEs, when paired with sequential coordinate selection, as a strong and promising foundation for vision interpretability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.