acceptodds
Under review as a conference paper at ICLR 2027

What Pooled Explained Variance Can Hide: Evaluating Crosscoder Fidelity in Vision Encoders

Abstract

Crosscoders, a variant of sparse autoencoders, were initially developed to compare representations across models and to analyze features across layers. However, their use for comparing blocks within a single vision encoder remains underexplored. We ask whether raw explained variance reflects typical-token reconstruction and predicts VLM behavior after a reconstruction is placed back into the encoder. We study this limitation using crosscoders, which reconstruct representations across layers and models. Our analysis shows that pooled explained variance weights each token by its squared distance from the activation mean, allowing a small high-norm population to dominate the score and obscure weaker reconstruction elsewhere. We evaluate crosscoders across five vision encoders using masked, token-balanced, and norm-stratified fidelity, separately accounting for population recentering. Controlled sparsity sweeps reveal that typical-token reconstruction deteriorates faster than raw explained variance as crosscoders become sparser. We then test these reconstructions inside four VLMs on the SpatiaLab spatial-reasoning benchmark. In Qwen2.5-VL, reconstructions explain over 99% of pooled activation variance yet reduce accuracy by up to 6.9 percentage points. Additional encoders and a separate stress test support the broader discrepancy between pooled fidelity and preserved behavior. Interventions identify model-specific multilayer perceptron contributions necessary for the measured high-norm contrasts, which later blocks often inherit, while Gemma provides a control without a persistent high-norm tail. Together, these findings establish the need to evaluate reconstruction across token populations and test its behavioral consequences. We recommend reporting raw explained variance alongside token-balanced and norm-stratified fidelity, followed by block replacement to assess whether reconstruction errors affect the behavior being studied.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.