Why Patch Selection Rarely Saves Compute, and How to Fix It
Abstract
GigaPixel Whole slide image (WSI) analysis is computationally expensive because pathology foundation models have to process large numbers of image patches extracted from WSI. Recent advances tend to select only a candidate patches for slide-level aggregation and representation. However, the selection process is usually carried on high-magnification resolution which requires heavy computation for the task processing. To avoid such computation burden, the process of patch selection/embedding can be utilized at low-magnification level of WSI; but coming with the price of losing pixel information. We address this problem with a three-stage coarse-to-fine cascaded process to progressively select patch regions from low- to high- magnification levels in WSI; imitating a very similar process of pathologists during microscopy diagnosis. A frozen thumbnail encoder first screens all candidate regions (i.e. global tissue view), followed by a frozen pathology foundation model that re-ranks the survivors using one coarse view per region. Only the final selected regions are then encoded at higher resolution. All encoders are frozen, and two lightweight linear attention scorers are trained with attention distillation to transfer the coarse ranking to the thumbnail stage. Across three classification and five survival datasets, we demonstrate the our cascade processing significantly reduces the inference GPU time and resource by –, while preserving the task performance evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.