Beyond Retraining Stability: Auditing Duplication Dependence in Vision Sparse Autoencoders
Abstract
Sparse autoencoders learn directions from a chosen activation corpus, making their features sensitive to both model representations and the distribution of training examples. Can a direction recur across retrainings yet depend on repeated exposure to particular images? We investigate this question through controlled near-exact duplication in activation corpora from frozen vision transformers. Our audit compares each reference direction’s displacement after corpus intervention with its variability across unchanged-data retrainings. We contrast full removal with copies-only removal, which retains the original image, and include size-matched placebos for both. Directions that are stable under unchanged-data retraining can lose their closest counterparts when repetitions are removed. Above the estimated detection threshold, copies-only removal reproduces most of the full-removal effect and identifies substantially overlapping directions. Wider dictionaries exhibit lower detection thresholds across independent paired experiments, although absolute thresholds vary. These thresholds remain audit-dependent, and calibration diagnostics reveal tail inflation among stable latents, precluding an unconditional false-discovery-rate guarantee. Furthermore, combinations of retrained features recover much of the reference activation signal, distinguishing dedicated-direction loss from information loss. Our findings show that retraining stability alone does not establish robustness to duplication in the SAE training corpus.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.