Perturb-Path: Benchmarking H&E Pathology Foundation Models with in vivo Genetic Perturbations in a Lung Cancer Model
Abstract
Pathology foundation models (PFMs) trained on datasets of histopathology images can learn robust patient representations that transfer to downstream disease modeling and drug response prediction tasks. While PFMs are trained on large amounts of unlabeled histopathology data, the scarcity of meaningful clinical labels forces these models to be evaluated on much smaller cohorts. Moreover, the labels in these evaluation cohorts are often noisy and corrupted by hard-to-resolve confounded factors. To mitigate these difficulties, we sought to create a PFM evaluation dataset with rich biological diversity, carefully-measured labels, and biological replicates. We leveraged Perturb-map to generate hundreds of genetically diverse tumors in a mouse model of lung cancer with paired hematoxylin & eosin (H&E) and multiplex immunofluorescence (mIF) imaging readouts. The resulting Perturb-Path dataset includes 84 knockout perturbations across 32 mice, resulting in 19,018 identified tumors with 116,383 H&E patches. We release a new PFM benchmark suite around this dataset and assessed eleven publicly-available PFMs and three ImageNet-based baselines. The benchmark suite evaluates the models on their ability to represent a range of phenotypic factors, including retrieval of genotypic perturbation labels and prediction of tumor microenvironment composition derived from the paired mIF readouts. This is, to our knowledge, the first publicly-available pathology benchmark dataset involving gene perturbations. We find that PFMs varied in their ability to retrieve specific perturbations, suggesting the presence of a genetically-associated phenotype that is captured by some models but not others. Further, we identify differences in ability to predict macrophage fraction versus T cell fraction across models, highlighting specific areas of improvement for model development. Overall, we found the PFMs outperform the ImageNet baselines, indicating that training on histopathology data of human origin translates to improvement on tasks in a mouse model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.