Per-spot semantic conditioning for H&E to spatial-transcriptomics generation
Abstract
Predicting spatial transcriptomics (ST) from hematoxylin and eosin (H&E) images is typically performed using image encoders coupled with regression models trained under mean-squared error (MSE). Under MSE, the optimal predictor is the conditional mean, whose variance is necessarily lower than that of the true conditional distribution. We show that this structurally limits the dynamic range and spatial structure that regression can recover, not because of insufficient model capacity but because of the learning objective itself. Across HEST-bench tasks, regression compresses the predicted dynamic range to 0.74× that of the ground truth on average, with the strongest compression on the hardest tasks, while Pearson correlation largely hides this failure by remaining insensitive to spatial over-smoothing. This limitation persists even in high-capacity models and can be characterized through an aleatoric uncertainty decomposition, suggesting that H&E-to-ST models should be evaluated on dynamic range and spatial structure alongside correlation. Motivated by this limitation, we reformulate H&E-to-ST prediction as conditional generation over the Visium spot grid. Our framework introduces user-specified per-spot semantic conditioning, enabling expression to be steered along biological axes that are not recoverable from morphology alone. On a morphology-invisible interferon-response axis (image predictability r=0.06), we achieve accurate directional control while evaluating only on genes disjoint from those defining the conditioning signal. A continuous per-spot scalar consistently outperforms a tertile one-hot representation. When the control axis is instead predicted from H&E, its correlation with the target remains low (r=0.13) and controllability collapses, demonstrating that the conditioning signal must be explicitly specified. Finally, we couple conditional generation with a regression anchor that preserves positional fidelity while allowing generation to restore biologically plausible spatial variation suppressed by regression. On held-out tissues, the proposed model matches regression-level localization while improving spatial-structure fidelity beyond what Pearson correlation captures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.