A Concept-Based Interpretability Benchmark for Pathology Foundation Models
Abstract
Pathology foundation models are usually compared by downstream accuracy, which can hide differences in what their representations encode and what downstream predictors use. Existing concept-based methods often use concepts or explanations that depend on the model being analysed, which makes direct cross-model comparison difficult. We introduce a model-independent framework that compares foundation models using the same predefined pathology concepts. We derive 23 morphology factors from structured pathology reporting standards and ground them in 9,672 caption-asserted factor–panel examples from 4,586 open-access articles. Using factor-specific probes on 16 frozen encoders, we measure factor accessibility, or how well each factor can be linearly recovered from a representation. Encoders differ in both overall accessibility and their factor-level profiles, and models with similar overall accessibility can differ substantially in which factors are most accessible. We then remove factor-associated information using the corresponding factor directions and measure factor sensitivity, or how much downstream performance changes after removal. In melanoma relapse prediction, accessibility is only weakly related to sensitivity. Across three melanoma prediction tasks, sensitivity patterns are consistent across retraining but differ across encoders and tasks. Our framework therefore provides a common concept space for comparing both what foundation models represent and what factor-associated information downstream predictors are sensitive to, turning concept-based interpretability into a tool for comparative foundation-model analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.