When Weak Features Hide: Separating Statistical Detection from Sparse Autoencoder Recovery
Abstract
To understand weak features missed by sparse autoencoders (SAEs), we distinguish three questions: whether a statistical test detects a feature, whether a decoder can represent it, and whether training recovers it. We analyze covariance tests that refit coordinate-wise variances and derive a closed-form visibility factor that vanishes for features aligned with the detector's coordinate axes. Such signals can be absorbed into the fitted variances, creating a test-specific blind spot that does not restrict an unconstrained SAE decoder's representational capacity. We also derive a held-out alignment certificate under rank-one Gaussian covariance and recovery guarantees for decoder families that explicitly preserve a data-estimated anchor. In a prespecified synthetic control, an isotropic-null-calibrated covariance statistic has power 0.081 for an axis-aligned feature and 1.000 for a dense feature, while matched TopK SAEs recover both in all eight seeds. A post-hoc correlation-based sensitivity analysis yields the same qualitative contrast. A separate post-hoc experiment holds the architecture and non-anchor initialization fixed: training-PCA anchors yield 183/192 recoveries versus 21/192 for random anchors. The same comparison yields no recovery-rate advantage on the tested injected GPT-2/Pythia residual backgrounds. These results separate detector-specific invisibility from decoder feasibility and clarify the assumptions required for certified recovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.