BEYOND STATIC BENCHMARKS: LIVING EVALUATION FOR SINGLE-CELL FOUNDATION MODELS
Abstract
Single-cell datasets and sequencing technologies are evolving rapidly, yet the evaluation of single-cell foundation models (scFMs) still relies largely on fixed benchmarks that may overlap with expanding pretraining corpora. We introduce scLiveEval, a living, contamination-aware framework that continuously incorporates newly released datasets and updates scFM evaluation over time while screening for potential data exposure using available temporal and provenance evidence. Study-defined novel cell populations are used to evaluate generalization to emerging biology; comparisons across platform-specific cohorts assess whether relative model rankings changes across measurement contexts; and, as newly annotated datasets enter the benchmark, human–model annotation disagreement analysis examines divergence between model predictions and supplied labels, highlighting populations that require further review and informing model applicability. We observe clear context dependence in model behavior: novel-population detectability varies across datasets, while biological preservation, batch integration, and model rankings differ across platform-specific cohorts. Cross-model consensus also reveals recurrent annotation disagreements supported by donor-specific transcriptional evidence. These findings motivate a shift from static benchmarking toward living evaluation as single-cell data and technologies continue to evolve.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.