Diagnosing spatial prediction in histology representations
Abstract
Predicting the cellular and molecular state of surrounding tissue offers a test of what histology representations encode beyond a local image. A spatial prediction score, however, can rise for reasons other than prediction of unseen tissue: the comparison can hide the baseline, the target can favor one encoder, and the input can see part of the target. We develop a diagnostic protocol that isolates each effect with matched evaluations across thirteen TCGA cancers and colorectal (Xenium) and renal (Visium) spatial-omics cohorts. Retrained TRIPLEX and MERGE show 74–92% smaller context gains on raw expression than on targets averaged over neighboring tissue already visible to the model, and neither method beats the training mean on raw renal targets. With six frozen encoders, an advantage in predicting nearby over distant tissue coexists with both predictions below the training mean, and the UNI2-H versus ResNet50 gap shrinks by 77–78% on shared measured targets while pathology encoders keep their ranking. Notably, images of the target region predict it where local crops cannot, and two slide encoders show no gain that clears a prespecified criterion once the target is withheld. Local UNI2-H features still improve nearby cell-composition prediction beyond a reference from the crop's measured gene profile, and a case study of a spatially supervised encoder shows that its neighborhood gain depends on the target, readout, and baseline. We distill these diagnostics into three reporting requirements: absolute scores for both compared targets, shared measured targets for encoder comparisons, and a declaration of the tissue each input and target can see.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.