Every Gene Count Has a Platform and a Slide: Measurement-Aware Observation Model for H&E-to-Spatial-Transcriptomics Prediction
Abstract
Predicting spatial transcriptomics (ST) from H&E images is moving toward foundation models trained by pooling data across cohorts and platforms, on the assumption that a larger pool makes the model more generalizable in a new cohort. We tested this assumption with five H&E-to-ST models and found that the larger pool gave no consistent gain on the HEST benchmark or across downstream tasks. The cause, we argue, lies in the target, not the model. We first clarify that gene counts, the prediction target, are measurements rather than expression itself, and show how far this shifts the target: expression is read at a scale that differs by platform and gene, at a level that varies from slide to slide, and as counts whose noise varies by gene and platform. The standard objective, squared error on log counts, has a term for none of these, so the output must account for everything the measurement contributed. We then propose an observation model that attaches to any regressor without changing its architecture, with a term for each of the three: the per-gene scale of each platform, estimated once and held fixed; the variation of each slide, absorbed in training by an intercept of its own; and the counts themselves, modeled with a negative binomial likelihood. Under our observation model, the same five models all improve with pooling on the benchmark. The improvement carries over to three further tasks: detection of tertiary lymphoid structures, spatial clustering, and evaluation on a new platform, with a gain over the standard objective in nearly every comparison. Our code is available at https://anonymous.4open.science/r/ST-Obs/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.