Iterative distillation of Matryoshka sparse autoencoders improves hierarchical feature organization in ESMFold2
Abstract
Matryoshka sparse autoencoders (SAEs) can decompose model representations into interpretable features, but reconstruction quality alone does not indicate how strongly those features influence model behavior. Furthermore, learned features can also vary across runs, complicating their characterization and reproducibility. These challenges are particularly difficult to study using large language models (LLMs), since SAE dictionary sizes capture only a fraction of an LLM’s feature space. In contrast, protein folding models may provide a better test case because protein structure is organized across multiple scales, from individual amino acids and secondary structure to domains and global structure. Here, we apply Distilled Matryoshka SAEs (DMSAEs) to study activations from ESMC-6B in ESMFold2. We assign each latent an attribution score using gradient activation with respect to ESMFold2’s predicted local distance difference test (pLDDT) and predicted template modeling (pTM) scores. These metrics reflect the model's confidence in the local and global structure, respectively. We find that distillation improves the recovery and hierarchical placement of features associated with individual amino acids across independent training runs. We also find that DMSAE distillation produces more reproducible and clearly separated rankings of amino acid features than existing SAE baselines. These results show that Distilled Matryoshka SAEs can refine the coarse ordering imposed by nested Matryoshka prefixes and improve the consistency of learned feature hierarchies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.