acceptodds
Under review as a conference paper at ICLR 2027

Split, Not Sparse: Ground-Truth Feature Superposition Across Scientific Domains and Model Scale

Abstract

Molecular and protein transformers predict properties such as hERG potassium-channel blockade and DNA binding from chemical structures and amino-acid sequences. Explaining these predictions requires identifying which scientific properties the models represent and how the associated directions affect their outputs. Sparse autoencoders (SAEs) express internal activations through learned features, with only a small subset active per input. Evaluating these features requires measuring both their coverage of a property across examples and prediction sensitivity to removing their associated directions. We introduce a paired recovery-and-removal protocol that ranks features against an external annotation, decodes it from feature groups, and removes the corresponding decoder directions with a random-removal control. We report three findings. (1) Group Recovery: groups improve property decoding in a ChemBERTa hERG predictor and three ESM-2 DNA-binding predictors. For the molecular conjunction of a basic nitrogen and an aromatic system, top-64 features reach 0.842 AUROC, versus 0.740 for one feature and 0.853 for the full layer. This recovery curve supports feature-budget selection. (2) Recovery–Intervention Dissociation: ESM-2 top-64 decoding stays within 0.026 AUROC of the full-layer reference, while targeted-minus-random removal effects range from 0.041 at 35M to 0.002 at 650M under the evaluated budgets. The paired measurements distinguish property recovery from removal sensitivity. (3) Target-Dependent Coverage: an aromatic-plus-basic protein target leaves reference gaps up to 0.116 AUROC at top-64, identifying settings where the selected group provides incomplete coverage. Together, the protocol and findings support choosing feature groups for a specified scientific property and evaluating their associated prediction effects.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.