acceptodds
Under review as a conference paper at ICLR 2027

OmicsBench: A Large-Scale Verifiable Benchmark for Agentic Life Sciences Capabilities

Abstract

Progress in the life sciences relies on the collection and analysis of complex data modalities. Increasingly capable AI agents have shown promise in revolutionizing the latter by automating bioinformatics. To realize this progress, the field has devised a number of agentic life sciences benchmarks. These benchmarks often test agent capabilities broadly, with a few tasks per field across many life sciences domains. Consequently, they do not exhaustively explore capabilities in a particular bioinformatics domain. We present OmicsBench, a collection of 784 expert-constructed, short- and long-horizon tasks that comprehensively explore six data-intensive life sciences domains (spatial-omics, single-cell transcriptomics, epigenomics, variant interpretation and discovery, microbial surveillance, and protein function). Each task is composed of a real-world dataset, a task prompt, and a deterministic grader. We show that across 31 model-harness pairs, no agent scores higher than 50% on the benchmark and conduct detailed analysis on failure modes and model behavior. It is our hope that our work will inform model capability development and future construction of benchmark and training tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.