Different Features, Reproducible Responses: Sparse Autoencoders Across Model Changes
Abstract
Sparse autoencoders (SAEs) represent model activations using sparse combinations of learned features. Each learned feature has a latent activation; features are often compared across dictionaries by matching these latents. We ask whether this is sufficient to assess the reproducibility of an observed activation response. Across independently fitted sparse dictionaries before and after model adaptation, compact combinations recover responses that no affine single-latent predictor recovers to the same accuracy. The singleton comparator searches the entire target dictionary and optimizes directly on evaluation inputs; group memberships and coefficients are fixed beforehand. We compare three ways to select groups: unbalanced optimal transport (UOT), descriptor cost, and calibration-response correlation. Across TinyLlama and Qwen3, with five fits per endpoint, compact groups increase recovery across all target fits. Token responses and unchanged-model refits show the same qualitative separation, and simple cost-selected groups outperform UOT groups. These findings distinguish the recoverability of an observation from the reappearance of an individual latent. Cross-dictionary references can retain compact realizations and their measured accuracy, allowing response evidence to accumulate without a canonical feature vocabulary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.