Beyond Reconstruction: Preserving and Decomposing Compositional Behavior with Multimodal Sparse Autoencoders
Abstract
Sparse autoencoders (SAEs) decompose VLM representations into interpretable features, but recovering objects, attributes, and relations does not establish that an SAE preserves how the model combines them. Reconstruction, sparsity, retrieval, and feature purity therefore provide limited evidence of compositional faithfulness. We identify two requirements for valid multimodal analysis: independently inferred image and text codes to prevent pair-identity leakage, and determinate scores for every selected feature. Without these safeguards, random dictionaries can yield near-perfect retrieval, while fixed product-TopK can select coordinates arbitrarily. We then train a multimodal SAE using Group-Hoyer sparsity, cross-modal decoding, and a variable-size positive-only paired auxiliary. Across matched CLIP, SigLIP, and SigLIP2 experiments, it improves reconstructed SugarCrepe accuracy over MGSAE by 3.7-5.1 percentage points while preserving retrieval geometry and automatically measured shared-concept coverage with 64-256 active features. Exact compositional-margin decompositions identify compact, cross-seed-recoverable circuits whose ablation and activation swapping affect held-out decisions more than energy-matched random controls. Ordered relation binding, however, remains distributed rather than localized to individual symbolic features. These results show that multimodal SAEs can preserve and causally expose compositional computations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.