scOLAR: Ontology-Anchored Open-Set Annotation of Single-Cell RNA-seq Data
Abstract
Reference-based single-cell annotation becomes an open-set problem when target data contain cell types absent from the reference. Existing approaches can reject unfamiliar cells, but the evidence used to distinguish known from unknown cells usually does not explicitly model cell-type hierarchy. This is limiting because an unseen fine-grained type may share broad transcriptional programs with a represented lineage, causing lineage similarity to support an incorrect known identity. Rejected cells must also be organized without knowing the number of unseen types. We present scOLAR, an ontology-anchored framework that separates fine-type evidence from broader-lineage affinity. scOLAR learns Cell Ontology-indexed prototypes with hierarchy-aware objectives, applies a reference-calibrated rule for selective label acceptance, groups rejected cells without a supplied novel-class count, and assigns broader ontology context post hoc. Across benchmarks, scOLAR achieves an AUROC of 0.973 for novelty detection, the highest published-protocol Overall accuracy on all five datasets, and the highest Novel accuracy on four. Under a separate common evaluator, scOLAR achieves a 7.2-percentage-point advantage in macro novel recovery together with higher known-cell accuracy. Matched ontology controls show stage-specific differences, with the largest observed changes in routing and post-rejection lineage context rather than novelty ranking or rejected-cell grouping. scOLAR turns open-set rejection into structured analysis of unresolved populations. Code is available in an anonymous repository.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.