Batch-Robust Cell Representations by Self-Distillation
Abstract
Zero-shot embeddings from masked-reconstruction single-cell foundation modelssuch as scGPT and Geneformer do not reliably beat HVG+PCA at batch integration.Self-distillation methods such as DINO, which optimise the embeddingspace directly, lead imaging benchmarks but have not yet shown clear gains onscRNA-seq. We train scRAPTOR (single-cell RAPid Teacher-student Optimisationof Representations), a 1.8M-parameter attention-free encoder, by prototypeself-distillation on a public corpus of 86 million cells from 1,065 human studies.Sinkhorn equipartition keeps the prototype head from collapsing, and a low-rankstudy term on the prototype logits, discarded after training, absorbs study identity,so the encoder needs no batch label at inference. Zero-shot, scRAPTOR beatstransductive HVG+PCA on scIB Total on the Pancreas, PBMC and DKD benchmarks,and leads every other zero-shot embedding we scored (scGPT, Geneformer,scJEPA, Concerto and two scVI models trained on our corpus) on Pancreas, PBMCand the three-dataset mean. Without Sinkhorn equipartition the head collapses to adozen prototypes and loses held-out-study accuracy; without the study term theembedding carries more study identity and scores lower on every benchmark. Werelease the weights of all three seeds, the cell-level split and the atlas build scripts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.