PathRet: Benchmarking Whole-Slide Image Retrieval Beyond Label Agreement
Abstract
Whole-slide image (WSI) retrieval returns prior cases that pathologists can directly compare, but conventional label-based evaluation measures diagnostic agreement with the query rather than which slides are actually returned. Replacing a retrieved slide with another assigned the same diagnostic relevance at the same rank leaves the evaluation unchanged. Such metrics can therefore measure whether the right kind of case was retrieved, but cannot distinguish which instances were selected within that diagnostic group. We compare seven retrieval systems from different methodological families and analyze queries for which a pair of systems each returns five slides with the query’s exact subtype. In these cases, label-based evaluation regards both systems as fully successful. Yet among 474 such queries for PRISM and Yottixel-K, 72.8% have no slide in common between the two returned top-five sets. Thus, retrievals that appear equally successful under diagnostic labels can present substantially different cases to the user. To address this limitation, we introduce PathRet, a benchmark of 4,249 WSIs from 13 cancer types that evaluates both label-defined clinical relevance and visual relatedness of the retrieved slides under a fixed retrieval protocol. Visual relatedness is measured by embedding WSI patches with a fixed pathology foundation encoder and comparing the query and retrieved slides at the slide level; an auxiliary measure compares their visual-phenotype distributions. When applied to real retrieval systems, the two views can lead to different system comparisons. PRISM scores higher than Prov-GigaPath on subtype and stage relevance, whereas Prov-GigaPath scores higher on both slide-level representation similarity and phenotype-composition similarity. This shows that a system retrieving more diagnostically relevant cases does not necessarily select cases that are more visually related to the query. WSI retrieval should therefore distinguish retrieving diagnostically appropriate cases from selecting visually related cases within those diagnostic groups, rather than evaluating retrieval through a single view alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.