acceptodds
Under review as a conference paper at ICLR 2027

A Concept-Based Interpretability Benchmark for Pathology Foundation Models

Abstract

Pathology foundation models are usually compared by downstream accuracy, which can hide differences in what their representations encode and what downstream predictors use. Existing concept-based methods often use concepts or explanations that depend on the model being analysed, which makes direct cross-model comparison difficult. We introduce a model-independent framework that compares foundation models using the same predefined pathology concepts. We derive 23 morphology factors from structured pathology reporting standards and ground them in 9,672 caption-asserted factor–panel examples from 4,586 open-access articles. Using factor-specific probes on 16 frozen encoders, we measure factor accessibility, or how well each factor can be linearly recovered from a representation. Encoders differ in both overall accessibility and their factor-level profiles, and models with similar overall accessibility can differ substantially in which factors are most accessible. We then remove factor-associated information using the corresponding factor directions and measure factor sensitivity, or how much downstream performance changes after removal. In melanoma relapse prediction, accessibility is only weakly related to sensitivity. Across three melanoma prediction tasks, sensitivity patterns are consistent across retraining but differ across encoders and tasks. Our framework therefore provides a common concept space for comparing both what foundation models represent and what factor-associated information downstream predictors are sensitive to, turning concept-based interpretability into a tool for comparative foundation-model analysis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.