acceptodds
Under review as a conference paper at ICLR 2027

Does the Route Matter? Functional Validation of Shared–Private SAEs

Abstract

Sparse autoencoders (SAEs) are commonly evaluated by reconstruction and feature descriptions, neither of which establishes that a learned decomposition matters to downstream computation. We introduce an operator-matched test of functional routing validity: correct and wrong private routes replace only the component represented by a training-only low-dimensional transform while preserving the activation's orthogonal residual. We instantiate the test with a shared–private SAE estimator based on principal-vector contrasts and group-level horseshoe shrinkage. On repository-derived prose and code, correct routing lowers held-out next-token negative log likelihood at three GPT-2 and three Pythia-70M layers. On independent text, the effect replicates at all three Pythia layers but GPT-2 intervals include zero. A 10-seed Pythia experiment removes the algebraically forced shared-subspace intersection and retains positive route effects on repository ( ) and independent text ( ); 50%-randomized routing is intermediate. The same test transfers to RxRx1 cellular embeddings, where correct HUVEC/RPE routing lowers perturbation-label loss by and improves top-1 accuracy by percentage points. Calibrated simulations return to zero under known nulls. Functional routing validity is thus measurable across modalities while remaining model- and corpus-dependent.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.