Structured Generative Modeling for Pixel-Level Multiplex Protein Prediction from H&E with ProDICE
Abstract
Multiplex spatial proteomics enables spatially resolved characterization of protein expression but remains costly and difficult to scale, whereas hematoxylin and eosin (H&E) imaging is routinely collected and widely available. Existing histology-to-spatial-omics methods predict gene expression at either the multi-cell spot level or the single-cell level, with single-cell prediction typically requiring cell segmentation to delineate cells before molecular profiles can be assigned. Pixel-level approaches avoid this requirement, but most produce deterministic predictions for individual protein channels rather than modeling the joint conditional distribution of the multiplex protein panel, limiting their ability to capture cross-protein dependencies. We introduce tein iffusion from mage-onditioned mbeddings), a conditional generative framework for pixel-level prediction of multiplex protein expression from H&E. performs diffusion by modeling the joint conditional protein distribution given histological features, enabling the simultaneous prediction of spatially organized protein maps while preserving cross-protein structure. We evaluate on held-out subjects from colorectal cancer and heterogeneous multi-tumor cohorts against four competing methods. achieves the best overall performance across Pearson and Spearman correlation, SSIM, Wasserstein distribution distance, and variance retention, while better preserving protein-protein correlation structure. We further demonstrate clinical relevance of predicted protein profiles in downstream prediction tasks on TCGA-CRC. These results demonstrate the potential of joint generative modeling for scalable virtual spatial proteomics. Code is available at https://anonymous.4open.science/r/prodice-iclr2027-review-25CD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.