SPORE: Semantic-Aware Global Post-Training for Reusable Single-Cell Representations
Abstract
Pretrained single-cell foundation models encode broad biological variation, yet their representations are not explicitly organized to make cell identity, biological context, and gene semantics accessible for downstream use. We introduce SPORE, a semantic-aware global post-training framework that adapts a pretrained single-cell foundation model once with heterogeneous biological supervision and yields a reusable checkpoint for deployment across studies and tasks. SPORE couples gene-semantic embedding alignment, ontology-aware cell supervision, and structured biological-context alignment to expose complementary biological semantics across scales. We evaluate the resulting checkpoint without target-specific backbone fine-tuning across complementary tests of downstream transfer, representation structure, and direct semantic inference. Across controlled evaluations, SPORE improves downstream transfer, makes biological information more accessible in frozen representations, and supports direct cell-type inference in new studies. These results support semantic-aware post-training as a means of organizing complementary biological semantics into a reusable single-cell representation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.