What Does Reconstruction Fidelity Measure under Distribution Shift?
Abstract
Reconstruction fidelity, reported as variance explained or fraction of variance unexplained, is used to judge learned representations, including when a model fitted on one dataset reconstructs another. Under such distribution shift, one score mixes three quantities: credit from the reference it is measured against, error from the reconstruction's mean being in the wrong place, and the structure around the target's own mean that the reconstruction captures. Reference dependence and the split of mean-referenced squared error into mean bias and centered error are classical; for reconstructions confined to a convex set, as in archetypal analysis, we show that the score is at most the fit available after one free translation (the translated fit) minus the squared distance from the target mean to the set, so geometry alone can force a negative score. In cross-corpus transfer of archetypal dictionaries (K = 32) over text embeddings, 37 of 48 settings score below the target mean, yet their fit around that mean exceeds that of randomly rotated hulls, though it nearly matches hulls of randomly chosen source points; transferred dictionaries reach 32–77% of the target-informed translated fit of dictionaries fitted on the target corpus, and location is the larger share of the transfer excess in all 48 settings, at each of K = 8, 16, 32 and 64. Representing the same records by their original text makes every cross-corpus score negative. For the Gemma Scope sparse autoencoders we evaluate, the same split shows that on natural data these SAEs sit far from the zero-score boundary, and that their decoder bias supplies substantial location correction when the codes are held fixed. Reporting the three quantities separately keeps a negative score from being read as absence of target-centered fit.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.