Toward a Shared Objective for Omics Foundation Models: The PRIMO Benchmark
Abstract
Omics foundation models are trained and promoted as tools for drug development, but they are rarely evaluated in this context. Existing benchmarks mostly operate at the level of the cell, the gene, or the batch, but the decisions that advance a drug program are made at the level of the patient the sample came from: which indication to pursue for a given target, whether a candidate should progress into the clinic, and which patients should be enrolled to observe a treatment effect. In addition, tasks and datasets are often preprocessed and built differently from one study to another, making results difficult to compare. We introduce PRIMO (Patient Representations in Multi-Omics), an open benchmark of sample-level tasks, with a first version focused on bulk and single-cell transcriptomics. We curate 63 tasks from 45 public cohorts on 30 diseases, most of them in four therapeutic areas, covering these decisions through five task categories: treatment outcome, clinical activity, prognosis, endotype and perturbation response, including cross-species and cross-indication transfer. Every task follows the same curation process: a clear drug-development question, a public cohort that can answer it, data cleaning, and tests ensuring the contained signal is linked to the biology rather than to technical artifacts. Under a common evaluation protocol with subject-grouped, label-stratified cross-validation, we evaluate 11 foundation models trained on bulk, single-cell or both modalities in the zero-shot setting: their weights stay frozen, and a shared linear probe reads their embeddings. All entries are rated against a raw-expression baseline, so that every rating reads directly as the value a representation adds over the raw data. Since foundation models are fine-tuned in practice, these zero-shot ratings are a lower bound on their value. Even in this setting, the strongest foundation models reach or exceed the raw-expression baseline, while most remain below it. We release PRIMO as a public leaderboard on curated tasks, with ratings per therapeutic area, giving model developers, drug developers and clinicians a shared reference to track progress and optimize for.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.