RNA Foundation Models Use Biology Their Explanations Mostly Miss
Abstract
Identifying which nucleotides drive an RNA model's prediction enables biological discovery beyond prediction alone. Attribution methods provide this insight by scoring each nucleotide's contribution. Their genomic validation has mostly been on task-specific convolutional networks (CNNs) trained from scratch, yet these methods are now applied to pretrained transformer RNA foundation models (FMs). Whether they expose what an RNA FM responds to has not been rigorously tested against labelled biological signal. We use in-silico mutagenesis (ISM), an exhaustive single-substitution scan of the model, to score against the same labels as every method, which we term grounding. ISM shows how much of the labelled signal each model demonstrably uses, a reference every method can be read against. Across six conventional attribution methods, a transformer-specific propagation method (AttnLRP), six RNA FMs (20M–650M parameters) and six RNA tasks, ISM substantially recovers the biological signals on four tasks, yet the conventional methods recover this signal only partially, and recovery varies substantially across FMs of near-identical task performance. We also show that RNA FMs' own substitution responses identify structural base pairs that the conventional methods do not consistently expose. These findings motivate validating each attribution method on the particular fine-tuned model before drawing biological conclusions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.