acceptodds
Under review as a conference paper at ICLR 2027

Do SAE Gains Transfer? Lexical Panels and Sparse Readouts

Abstract

Sparse autoencoders (SAEs) are central to mechanistic interpretability, commonly evaluated through two primary lenses: activation reconstruction fidelity and downstream sparse readouts. However, conventional evaluations frequently rely on two unverified premises: that architectural advantages identified on a specific lexical panel generalize across independent evaluation sets, and that optimizing reconstruction fidelity directly enhances downstream semantic readouts. In this work, we present a systematic evaluation framework that rigorously investigates these assumptions across lexical panels, optimization objectives, and detector adaptation interfaces on Pythia-160M. First, by evaluating paired checkpoints across an audited panel of 120 previously unseen lexical members across 1,200 distinct articles, we show that measured architectural margins are highly sensitive to panel composition and evaluation sample budgets, demonstrating the critical need for multi-panel auditing in representation benchmarks. Second, through controlled fixed-support coefficient optimization on official-class SAEs, we achieve substantial reconstruction error reductions of 8.23% for BatchTopK and 5.48% for Matryoshka, while discovering that reconstruction gains fundamentally decouple from downstream sparse readouts. We establish both a constructive mathematical proof and continuous intervention trajectories showing that reconstruction improvements can invert linear rank orderings. Third, we reveal that detector adaptation protocols (frozen, refitted, or reselected) systematically govern observed readout gains, exhibiting divergent behaviors across architecture families. Finally, we prove that contribution-based top- gating ensures exact scale invariance under dictionary reparameterization. Our findings provide actionable methodological principles and a reproducible benchmark harness for faithful representation evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.