When Do Biological Reasoning Models Use Their Biological Inputs?
Abstract
Biological reasoning models use post-training to connect an LLMs to biological inputs, typically representations from a biological foundation model and biological text such as gene and pathway descriptions, functional annotations, and gene lists. Their benchmark accuracy is taken as evidence that the LLM reasons over their biological inputs. We test this assumption in six biological reasoning models across DNA, protein, and single cell tasks. We perturb one biological input while holding the query and other inputs fixed, construct evidence conflicts that pair the foundation model representation of one genome, protein, or cell with the text of another, fit linear probes to the representations the language model receives, and analyze reasoning traces against the biological inputs. We find that current post-training strategies do not ensure that foundation model representations contribute to task performance. Evo2 and ESM3 contribute little to BioReason and BioReason-Pro performance on the evaluated tasks. Shuffling the DNA sequence barely changes BioReason disease prediction accuracy, and in evidence conflicts the two models follow the text in 97.9% and 99.7% of cases. Linear probes trained on the Evo2 and ESM3 representations predict the task targets, so these foundation models encode information relevant to the task, but provide limited overall performance improvement to BioReason and BioReason-Pro. In contrast, foundation model inputs contribute to ChatNT, Prot2Text-V2, and CellWhisperer performance, and differentially expressed genes in the gene sentence contribute to Cell2Sentence-Scale performance. Across SFT and RL checkpoints of BioReason-Pro and 42 BioReason checkpoints, increases in accuracy do not imply greater performance contributions from biological inputs. In reasoning traces, BioReason traces misstate nucleotide changes, while BioReason-Pro traces describe functions omitted from final predictions under evidence conflicts. Future work should test whether post-training objectives that reward correct use of biological inputs improve their contribution to task performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.