Dictionary Size and Dead Latents Handling Jointly Set the Recovery-Optimal Sparsity of Top-K Sparse Autoencoders
Abstract
Sparse autoencoders (SAEs) are trained to reconstruct activations under a sparsity constraint, but what interpretability requires is the recovery of the features underlying them. In a Top- SAE toy model with a known ground-truth dictionary, we judge an SAE by the atoms it actually uses rather than by its nominal width, and partition its live atoms against the ground truth and against a second training run. Against the ground truth, the recovery-optimal sparsity is set jointly by dictionary size and dead-latent handling: below the ground-truth dictionary size it rises with the ratio of dictionary size to feature count, whereas above it the dead-latent handling sets its level, and when dead latents are appropriately handled, every feature at the optimum is carried by a single live atom. Against a second run, the same structure appears at the optimum, agreement and recovery come apart elsewhere—a low shared fraction arises even when every atom is correct—and the overlap between the runs, read as a capture–recapture experiment, estimates the number of features to within its sampling error at every feature count and dictionary size we tested.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.