Stack-JEPA: Molecular Graph Distillation for 3D Spatial Transcriptomics
Abstract
Spatial transcriptomics (ST) observes isolated two-dimensional sections, cutting anatomy and molecular neighborhoods at one arbitrary plane. Three-dimensional (3D) ST lifts that limit, but profiling a section by ST costs far more than imaging it by histology, so volumes are built by imaging every section with hematoxylin and eosin (H&E), profiling a sparse subset, and imputing the rest. Existing imputers are supervised and count-space: they regress the raw counts of one fixed gene panel from paired sections, so the objective is dominated by dropout entries, a new panel forces a refit, and accuracy is capped by an inter-section alignment computed once before the model runs. We introduce STACK-JEPA, a masked molecular graph-distillation framework for learning molecularly informed representations throughout a serial histology stack. On a profiled section, a frozen molecular teacher defines both per-location latent states and their local molecular relations. A stack student is trained to recover the held-out section’s molecular representation and relation graph from a profiled context section, the target H&E image, in-plane coordinates, and tissue depth. Frozen TERRA and UNI2-h encoders feed a transformer with 3D rotary position embeddings, while node- and relation-level distillation transfer molecular structure without placing the genepanel dimension in the backbone. After pretraining, the backbone is frozen and only a lightweight panel-specific count decoder is fitted. The same learned descriptors can replace raw expression in partial fused Gromov-Wasserstein transport, including when the target section has no measured transcriptome. Across multipe tasks, imputation on public 3D ST benchmarks and alignment under the PASTE2 subslice protocol, STACK-JEPA outperforms the supervised and generative state of the art, on imputation in both per-cell and per-gene correlation, and it aligns sections for which no transcriptome was measured, which expression-based aligners cannot handle at all.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.