Post-hoc Robustification of Pathology Foundation Models with a Vision-Language Judge
Abstract
Pathology foundation models are meant to describe tissue, yet their embeddings also record where and how a slide was acquired. Institution, scanner and staining leave batch effects strong enough that tiles can group by domain instead of by biology, which weakens transfer to new sites. Standard mitigation methods require large amounts of labeled training data, while label-free methods often struggle to remove domain-specific bias while preserving genuine biological signal. Multimodal large language models (MLLMs) can make robust judgements on biological similarity in images, but their utility in computational pathology is under-explored. To address the robustness issue in pathology foundation models, we propose RObustify PAthology models with a vision-Language judge (ROPAL), a label-free protocol in which MLLMs supervise their adaptation toward biological similarity. ROPAL asks an MLLM which tissue tiles are biologically alike across institutions and uses these judgements to fine-tune the foundation model, so that its embeddings group tissue by biology rather than by the domain it came from. We demonstrate the effectiveness of ROPAL through extensive empirical analysis across a diverse range of datasets, diseases, tasks, base models and robustification baselines. We release the code, training corpus and fine-tuned checkpoints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.