Does the Route Matter? Functional Validation of Shared–Private SAEs
Abstract
Sparse autoencoders (SAEs) are commonly evaluated by reconstruction and feature descriptions, neither of which establishes that a learned decomposition matters to downstream computation. We introduce an operator-matched test of functional routing validity: correct and wrong private routes replace only the component represented by a training-only low-dimensional transform while preserving the activation's orthogonal residual. We instantiate the test with a shared–private SAE estimator based on principal-vector contrasts and group-level horseshoe shrinkage. On repository-derived prose and code, correct routing lowers held-out next-token negative log likelihood at three GPT-2 and three Pythia-70M layers. On independent text, the effect replicates at all three Pythia layers but GPT-2 intervals include zero. A 10-seed Pythia experiment removes the algebraically forced shared-subspace intersection and retains positive route effects on repository ( ) and independent text ( ); 50%-randomized routing is intermediate. The same test transfers to RxRx1 cellular embeddings, where correct HUVEC/RPE routing lowers perturbation-label loss by and improves top-1 accuracy by percentage points. Calibrated simulations return to zero under known nulls. Functional routing validity is thus measurable across modalities while remaining model- and corpus-dependent.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.