One Query Is Enough: Adapting Frozen Multiple Instance Models
Abstract
Whole-slide image (WSI) classification increasingly relies on strong pretrained pathology encoders, yet the downstream multiple-instance learning (MIL) stage is still commonly adapted with a fully trainable attention network or transformer aggregator. This raises a basic question: how much task-specific adaptation is actually required once tile representations are strong? We study an intentionally minimal regime in which the tile encoder is frozen and the bag representation is adapted by a single learnable query. Given frozen tile embeddings, the query exponentially tilts their empirical distribution and forms a weighted expectation that is passed to a linear classifier. The representation-adaptation budget is therefore only one -dimensional vector beyond the task head. We show that this mechanism admits a precise characterization: query attention solves an entropy-regularized instance-selection problem, its representation Jacobian is a tile-feature covariance operator, and its temperature continuously interpolates between mean pooling and hard instance selection. The empirical study is designed around sufficiency rather than universal SOTA: the provisional result scaffold places one-query adaptation within one percentage point of a substantially richer full-data aggregator on fine-grained EBRAINS, while making its advantage larger in few-shot and external-cohort settings. Across three pathology encoders, the same scaffold predicts a narrow – point gap to the strongest in-domain aggregator, and the multi-query diagnostic saturates after one or two queries. These patterns motivate a broader hypothesis: with sufficiently strong pretrained pathology representations, much of downstream MIL adaptation may reside in how the bag is queried, rather than in relearning the representation itself.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.