acceptodds
Under review as a conference paper at ICLR 2027

Mechanistic Interpretability Findings That Generalize: Chaperone-Dependent Folding in ESMFold2

Abstract

Interpretability findings in scientific models are often supported by post hoc plausibility rather than held-out tests. We introduce a workflow that turns interpretability outputs into statistically testable predictions: feature terms discovered in examples sharing a biological property should have stronger effects in unseen members than in controls. We apply this workflow to an SAE over ESMFold2 representations and to GroEL/GroES clients: proteins that require this chaperone system for folding in vivo. Perturbations in 25 discovery clients identify 67 stable SAE terms. Their ablations have larger effects in eight held-out clients than in 19 matched controls for 56 out of 67 terms (), and their selected-to-random effect ratio is % higher in clients than in controls. When multiple selected terms are ablated together, client-control separation increases from approximately 3.8- to 7.3-fold across a set of randomly subsampled terms. After these tests, interpretations of the selected terms connect to known client properties and motivate new hypotheses about client-cavity charge patterning, compensatory core packing, and delayed interface closure. Our contribution is a cohort-conditioned discovery-and-evaluation workflow that turns latent feature interactions into held-out tests of biological selectivity. The framework accommodates multiple discovery methods and uses the selected interventions to nominate biological hypotheses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.