Predictive Summarisation in Evo 2
Abstract
Genomic causal language models learn representations through next-nucleotide prediction, yet these representations are also used to infer biological properties of DNA. We investigate how biological readout and predictive function diverge across depth in Evo 2 7B and 40B. Both models exhibit an abrupt representational transition at which variation across genomic positions concentrates in fewer directions. Across this transition, the decline in genomic annotation readout steepens while next-base probe performance improves, including when both are evaluated at the same positions. Restricting standardised transition states to their 16 leading principal directions and continuing pretrained computation preserves native next-base prediction within prespecified NLL and output-KL tolerances, while reducing annotation AUROC by in 7B and in 40B relative to intact transition states. The same restriction before the transition fails to preserve prediction. Annotation decline also persists across probe controls and among positions with similar predictive distributions. We interpret these findings as : a reorganisation of contextual representations for next-base prediction that reduces the accessibility of genomic annotations. Earlier representations selected during development outperform final embeddings by in mean annotation AUROC in a retrospective chromosome 19 evaluation, with benefits extending to gene-structure and conservation readouts. Some variant-difference readouts instead favour final embeddings, showing that layer preference depends on the biological target and feature construction. Next-nucleotide prediction provides indirect supervision for biological properties; representations organised for this objective need not be the most useful for biological analysis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.