acceptodds
Under review as a conference paper at ICLR 2027

Representation Redistribution in Sparse Autoencoders

Abstract

Sparse autoencoders (SAEs) decompose language model activations into sparse features, but related prompts can redistribute activity across these features even when the underlying request is preserved. Tracking isolated features can miss such distributed responses, while aggregate representation distances do not identify which changes recur or whether they matter downstream. We introduce the Redistribution Sensitivity Score (RSS), which identifies large feature responses that recur in a consistent direction, and evaluate whether the resulting supports persist on variants excluded from selection. Across 27 pretrained SAEs spanning eight language models, controlled prompt transformations produce a reproducible redistribution profile that persists on held out variants. We then connect these recurring representation changes to fixed linear readouts. An exact contribution factorization separates response magnitude, readout weight, directional alignment, and clean margin, showing why persistent redistribution need not imply decision failure. In two frozen safety readouts, combining RSS support erosion with the full clean margin yields recurrent failure AUROCs of 0.935 and 0.958 using median supports of 20 and 19 features, compared with 0.965 and 0.964 for the full readouts. These results provide a compact way to identify reproducible sparse representation changes and quantify their downstream consequences without equating representation change with failure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.