Benchmarking the Capabilities Needed to Design DNA
Abstract
Genomic foundation models are increasingly used to propose DNA sequences and edits. But successful design requires capabilities that current benchmarks do not measure. Here we introduce a benchmark that assesses DNA Design Readiness Across Functional Tasks (DRAFT), which includes sequence priors, functional prediction, biological context, genetic interactions, and selection for a desired response. Across eight tasks built on five experimental datasets, we find that model choice and how model representations are accessed both affect performance. Frozen-representation outputs can match or exceed models' native outputs, and functional post-training substantially improves their predictive capabilities. These outputs can also transfer across experimental assays and genes, suggesting that models capture functional patterns beyond the data used to fit individual predictors. At the same time, strong prediction of average effects can conceal weaker prediction of context-specific effects, and accurate prediction of combined mutation effects does not imply accurate prediction of their interactions. In design tasks, model predictions can reduce the number of candidates needed to find a desired response, but experimental selection reveals substantial room for better choices, particularly for context-specific and multi-cell objectives. Overall, DRAFT shows that genomic models contain useful information for DNA design, while highlighting biological context and genetic interactions as important remaining limitations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.