acceptodds
Under review as a conference paper at ICLR 2027

Executable annotations of SAE features

Abstract

Sparse autoencoder (SAE) features are used in practice through their descriptions. While features fire at a prevalence of , LLM descriptions are evaluated at 1:1. Rescored at natural prevalence, descriptions collapse to 0.01-0.05 and describe concepts that are 40-200 more prevalent than the feature. In this work, we show that annotating a feature is equivalent to learning a Boolean function over the one-hot encoding of a token and its context. We annotate every feature of three GemmaScope dictionaries by learning this function using a monotone linear threshold function (LTF) over (offset, token) literals, and evaluate it directly as a per-token detector on held-out text. These annotations outscore LLM descriptions on the field's own benchmarks, and score 5-24 higher at natural prevalence. Scored at this scale, the dictionary resolves into strata that balanced evaluation cannot distinguish, and the LTF weights expose multilingual concepts, boilerplate templates, and superposition directly. Applied without modification to a protein language model, the same method recovers established structural motifs such as zinc-finger spacing, Kelch repeats, and P-loop motifs from sequence-only representations. All 147k features across the three Gemma dictionaries are annotated and evaluated within 7 hours on a single quad-GH200 node, with no LLM in the pipeline. All annotations and code will be released publicly.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.