Mechanistically Understanding How a Biological Language Model Predicts Beneficial Mutations
Abstract
Language models have shown impressive results when applied to biological tasks. Concurrently, recent advancements in interpretability have produced methods that enable the computation of language models to be understood mechanistically. This paper examines what language models learn when trained on proteins. Specifically, we analyze CovFit, protein language model introduced by Ito et al. [2025] trained to predict viral fitness from input SARS-CoV-2 sequences. We apply interpretability methods to (i) identify a MLP "look-up" computation that generalizes across several SARS-CoV-2 mutations, (ii) identify a particular biological circuit for the L455F mutation and (iii) show that highly conserved sites in SARS-CoV-2 are used as "attention sinks" over the course of computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.