HiLO: Hierarchical Label-Aware Optimal Transport for Cross-Domain Alignment of Pathology Foundation Model Embeddings
Abstract
Pathology foundation models produce strong patch embeddings for H&E whole-slide images (WSIs). However, their embeddings remain highly predictive of non-biological factors such as source centre, scanner, and staining protocol. Consequently, downstream models may fail to rely on biologically meaningful features, which limits their clinical applicability. Optimal transport (OT) offers a principled way to align feature distributions across domains. However, standard OT algorithms exhibit quadratic memory and time complexity in the number of points, making them highly computationally expensive for WSIs. We propose HiLO, a hierarchical label-aware optimal transport framework for cross-domain feature alignment that can scale efficiently to large datasets. Building on hierarchical refinement, we compute globally optimal one-to-one correspondences between foundation-model patch embeddings from different source domains in linear memory. We embed label information into the transport cost by augmenting features with a label embedding. This preserves the low-rank cost factorization required for scalability and suppresses cross-class matches under label shift. The resulting one-to-one correspondences define a Monge map that transports patch embeddings from each source domain to the target domain, on which any downstream model can be trained. We evaluate our method on two publicly available pathology WSI datasets AGGC2022 and BEETLE. All experiments follow a leave-one-domain-out protocol, where each held-out domain is a scanner or centre depending on the dataset, and we report both in-domain (ID) and out-of-domain (OOD) performance. Compared with state-of-the-art OT, stain normalisation, and domain generalisation methods, HiLO achieves higher performance, while also attaining lower total transport cost than the OT baselines. In conclusion, HiLO is a scalable and flexible framework for cross-domain feature alignment that is applicable to large-scale pathology WSIs and can be extended to a wide range of downstream applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.