acceptodds
Under review as a conference paper at ICLR 2027

Beyond Retraining Stability: Auditing Duplication Dependence in Vision Sparse Autoencoders

Abstract

Sparse autoencoders learn directions from a chosen activation corpus, making their features sensitive to both model representations and the distribution of training examples. Can a direction recur across retrainings yet depend on repeated exposure to particular images? We investigate this question through controlled near-exact duplication in activation corpora from frozen vision transformers. Our audit compares each reference direction’s displacement after corpus intervention with its variability across unchanged-data retrainings. We contrast full removal with copies-only removal, which retains the original image, and include size-matched placebos for both. Directions that are stable under unchanged-data retraining can lose their closest counterparts when repetitions are removed. Above the estimated detection threshold, copies-only removal reproduces most of the full-removal effect and identifies substantially overlapping directions. Wider dictionaries exhibit lower detection thresholds across independent paired experiments, although absolute thresholds vary. These thresholds remain audit-dependent, and calibration diagnostics reveal tail inflation among stable latents, precluding an unconditional false-discovery-rate guarantee. Furthermore, combinations of retrained features recover much of the reference activation signal, distinguishing dedicated-direction loss from information loss. Our findings show that retraining stability alone does not establish robustness to duplication in the SAE training corpus.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.