acceptodds
Under review as a conference paper at ICLR 2027

LineageSAE: Versioned Feature Vocabularies across SAE Retraining and Model Updates

Abstract

A feature index in a sparse autoencoder (SAE) is useful only if it keeps denoting the same feature after the SAE is retrained or the language model changes. Existing methods either encourage reproducibility or align independently learned dictionaries after training. Neither decides, for a saved parent and a changed view, which of its addresses to retain and which to free for new features. We introduce LineageSAE, a parent–child SAE that builds a versioned feature vocabulary from a saved parent and current activations without replaying the parent's data. A corrected marginal-utility test marks each parent address as inherited or released, and a parent-guided decoder handoff transfers reconstruction to the child. A temporary ranking loss uses the parent's pre-activation scores to align released addresses across independently trained children, including when ReLU outputs are zero. On SynthSAEBench and toy superposition with feature deletion, addition, and frequency changes, LineageSAE reaches child–child support Jaccard, the overlap of firing samples at the same address, of and without post-training assignment. The strongest matched baselines reach and . Reconstruction is close to -regularized TopK on SynthSAEBench; on Toy the NMSE is , against and for RA-SAE and -regularized TopK. A trained child keeps its index assignments and can parent the next update.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.