Cell-level representations in pathology foundation models are a question of read-out
Abstract
Pathology foundation models (PFMs) are trained and evaluated on tiles or slides, yet many biologically relevant tasks are defined at the level of individual cells. A cell's identity depends both on its own morphology and on its surroundings, so a cell embedding should be specific to the cell while integrating its context. Current read-outs trade one for the other. One approach crops a small window around the cell, resizes it to the encoder's input size, and keeps the CLS token. This isolates the cell, but it discards the surrounding tissue and presents the image at a magnification far from the one used in pretraining. Keeping the CLS token of a wider tile retains context, but the embedding summarizes the whole tile and is shared by many, often heterogeneous, cells. We study spatial indexing, which removes this trade-off: a tile is centered on the cell, and the patch tokens at its location are read out. The tile size sets how much context the encoder integrates, the token position sets which cell is represented, and DINOv2's token-level iBOT objective makes these tokens informative. To evaluate read-outs at scale, we curate 111 public Xenium samples with registered H&E, comprising 20.75 million cells with gene expression and expression-derived cell-type labels, the largest cell-level annotated pathology dataset to date. Across four PFMs, spatially indexed tokens outperform the CLS token of the same tile by 0.10 macro-F1 on cell-type classification, crop-and-resize is the weakest read-out, and the gains extend to single-cell gene-expression prediction. Linear probes and iBOT-token masking confirm that a single token holds both sides of the trade-off and that both pieces of information are complementary for classification. Spatial indexing is also cheaper: reading out every cell of a tile from a single forward pass needs 1.4% of the forward passes at a small cost in accuracy, and because the tokens already carry context, the neighbor-aggregation module of a gene-expression model can be removed, cutting its compute 26. We release the resource to support the development of cell-level pathology models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.