acceptodds
Under review as a conference paper at ICLR 2027

Don't Mind the Gap: Interpretable Vision-Language Models via Independent Sparse Autoencoders

Abstract

Contrastive Vision-Language Models (VLMs) produce general-purpose, multi-modal representations, but their black-box nature hinders safe and ethical deployment. Sparse Autoencoders (SAEs) can disentangle these dense representations into interpretable features, but the application of a single SAE across modalities creates a split-dictionary phenomenon, where learned features activate for a single modality, preventing the discovery of shared features. Current approaches attempt to bridge this gap using a co-activation loss, but we find that forcing a single shared SAE inherently compromises reconstruction quality, individual modality interpretability, and zero-shot performance. Instead, we propose decoupling the SAEs by modality while retaining the co-activation guidance. This dual architecture discovers features that retain their modality-specific nature while being synchronized to enable consistent cross-modal mapping. Our approach achieves superior interpretability and downstream zero-shot performance, especially in challenging high-to-medium sparsity regimes. Our findings indicate that exploiting the modality gap, rather than collapsing it, is the key to interpreting contrastive VLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.