ViO-Align: A Decoupled Vision-Omics Alignment Method for In Situ Cell Phenotyping
Abstract
Routine hematoxylin and eosin (H&E) slides preserve both cellular morphology and surrounding tissue architecture, providing a practical basis for phenotyping cells in their native microenvironment. However, conventional vision models are mainly trained with morphology-defined labels, while molecularly supervised approaches often operate on cropped or prelocalized cells and therefore do not directly connect molecular supervision with raw-patch in situ phenotyping. ViO-Align is introduced as a decoupled vision–omics framework that transfers single-cell molecular supervision to H&E-only in situ cell phenotyping. In Stage I, matched H&E cell crops and transcriptomic profiles are aligned through morphology-conditioned gene gating to learn molecularly guided single-cell representations. In Stage II, cells are localized from uncropped H&E patches using count-guided flow, while center-masked microenvironment aggregation captures complementary local tissue context. Cross-stage alignment transfers Stage-I molecular supervision to raw-patch representations, and downstream phenotyping explicitly combines target-cell morphology with the corresponding microenvironment representation. Only H&E images are required at inference. Across four established cell phenotyping benchmarks, ViO-Align improves macro-F1 by 4.67 absolute points on average over the strongest strictly test-disjoint baseline on each dataset, while maintaining competitive cell localization performance. To more directly evaluate molecularly guided representation learning, five molecularly defined cell-state recognition tasks are further considered. ViO-Align achieves the highest AUROC on all five tasks and improves over the strongest vision baseline by up to 3.86 absolute AUROC points. The learned H&E representations also transfer to patient-disjoint breast cancer diagnosis. Overall, ViO-Align achieves strong and consistent performance across cell phenotyping, molecular-state recognition, and diagnostic transfer, demonstrating robust generalization across cellular and tissue-level tasks. Code is provided in the supplementary material to facilitate reproducibility.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.