acceptodds
Under review as a conference paper at ICLR 2027

Persistent Sparse Autoencoders: Learning Feature-Specific Timescales in Language Model Representations

Abstract

Sparse autoencoders (SAEs) decompose language model activations into sparse features, yet these models traditionally encode each token independently, failing to expose information that persists across a sequence. We first show that temporal persistence can naturally emerge in standard SAE features: after a feature activates, the hidden state remains aligned with its direction, and past activations reduce the reconstruction error of later hidden states. How long this lasts varies widely across features. We therefore introduce Persistent Sparse Autoencoders (Persistent SAEs), an extension of standard SAEs that learns a persistence coefficient for each feature, allowing feature-specific timescales to emerge from reconstruction. Our experiments show that Persistent SAEs retain competitive reconstruction quality while learning a spectrum of timescales: short-timescale (fast) features remain locally interpretable, whereas long-timescale (slow) features accumulate higher-level contextual information. Moreover, we show in a prompt-injection monitoring case study that slow features retain injection signals over long contexts and causally affect monitor judgements. These results suggest that Persistent SAEs offer new opportunities for interpreting and monitoring language models via persistent features.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.