PixStain: Pixel-Space Diffusion with Cross-View Representation Alignment for Virtual Staining
Abstract
Virtual staining translates digital histology images between staining modalities, offering a faster and less expensive alternative to immunohistochemistry (IHC) and multiplex immunofluorescence (mIF). Existing diffusion-based virtual staining approaches generate high-quality images, but they operate in a VAE latent space that discards fine histological detail, require iterative sampling, and rely on output-space supervision alone. We observe that pathology foundation model (PFM) features provide informative spatial priors for virtual staining: their similarity maps closely align with regions of high marker expression, whereas those derived from VAE latents are dominated by noise. We therefore propose PixStain, a pixel-space diffusion model for virtual staining built on cross-view representation alignment. PixStain conditions a pixel-space Vision Transformer on multi-level PFM features of the H&E image and trains it with flow matching under x-prediction, so that a single stage of training yields the target marker image in a single inference step. Cross-view alignment matches shallow student features from spatially disjoint marker noising and H&E masking to deeper teacher features from a cleaner view. These complementary corruptions encourage the student to integrate information from both inputs. On three publicly available same-section datasets covering H&E-to-IHC, H&E-to-mIHC, and H&E-to-mIF, PixStain achieves state-of-the-art performance across pixel-level and perceptual metrics, as well as cell classification and segmentation metrics, except for pixel-level metrics on HEMIT, where it ranks second.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.