acceptodds
Under review as a conference paper at ICLR 2027

Mechanistically Understanding How a Biological Language Model Predicts Beneficial Mutations

Abstract

Language models have shown impressive results when applied to biological tasks. Concurrently, recent advancements in interpretability have produced methods that enable the computation of language models to be understood mechanistically. This paper examines what language models learn when trained on proteins. Specifically, we analyze CovFit, protein language model introduced by Ito et al. [2025] trained to predict viral fitness from input SARS-CoV-2 sequences. We apply interpretability methods to (i) identify a MLP "look-up" computation that generalizes across several SARS-CoV-2 mutations, (ii) identify a particular biological circuit for the L455F mutation and (iii) show that highly conserved sites in SARS-CoV-2 are used as "attention sinks" over the course of computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.