acceptodds
Under review as a conference paper at ICLR 2027

DIAL M FOR MONOSEMANTICITY : RECONSTRUCTION METRICS CANNOT CERTIFY SAE BASIS QUALITY

Abstract

Reconstruction metrics are the most common way to certify a sparse autoencoder (SAE), yet any functional of the reconstruction, MSE, FVU, CE-recovered and patched task accuracy among them, is provably constant under an invertible remix of the code: it cannot tell a dictionary from one strictly more mixed in the tar- geted latent basis. We use this identity constructively, mixing the basis of one pretrained SAE post hoc with the reconstruction preserved exactly, against a clone- split control arm whose residual support confound is a step rather than a ramp and so cannot manufacture a dose-response. Reading the result requires a calibrated operating point: the identical grid returns opposite signs depending only on the steering coefficient, and at the strength our pre-registration inherited the model sits 17–31 nats of KL from unsteered behaviour. At a post-hoc strength inside the regime our diagnostics certify, increased mixing lowers top-k SAE-latent steer- ing while every reconstruction metric reports the dictionary unchanged: −0.298 [−0.509, −0.089] on irony, negative but unresolved on SNLI, and negative in 21 of 23 cells outside the broken regime. Full-set steering is provably unaffected. We recommend steering results report answer-position KL, an off-target accuracy delta, and one calibrated strength as a fraction of the activation norm.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.