Quantile-Domain Multiple Instance Learning for Whole-Slide Image Classification
Abstract
Whole-slide images (WSIs) contain billions of pixels, making end-to-end modeling costly in computation and storage. Their classification is therefore typically formulated as Multiple Instance Learning (MIL) over patch features. However, some methods construct bag-level context with positional structures, neighborhoods, or approximations that depend on patch indices. Their predictions can therefore change with patch order. We show that patch interactions need not cause this sensitivity and that it can instead arise from index-addressed operators. Based on this insight, we introduce Quantile-Domain Multiple Instance Learning (QDMIL), a three-stage framework. First, it forms a fixed-size, permutation-invariant quantile grid over multiple learnable projections. This grid captures the slide-wide distribution of morphological features. Second, QDMIL uses the grid to derive soft quantile coordinates for each patch and bag-shared affine modulation. Together, these signals contextualize the patch representations. Finally, QDMIL fuses attention-pooled local discriminative evidence with a separate grid encoding for slide classification. We establish sufficient operator-level conditions for permutation invariance and prove that QDMIL satisfies them under deterministic inference and exact arithmetic. Against 15 baselines across four datasets, QDMIL achieves the highest mean accuracy and F1-Score in all eight dataset–metric comparisons. Paired order-control experiments across seven methods further support distribution-relative patch positioning as a general principle for permutation-invariant bag-level context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.