SlideTTT: Slide-Specific Test-Time Training for Generalizable Whole-Slide Images Analysis
Abstract
Multiple instance learning (MIL) has become the dominant paradigm for weakly supervised whole-slide image (WSI) analysis, where large-scale WSIs are divided into patches and their representations are subsequently aggregated for slide-level prediction. However, most existing MIL methods still follow a static inference paradigm, applying a fixed aggregator learned during training to all test WSIs. In the presence of substantial inter-slide heterogeneity and cross-center distribution shifts, such a unified aggregation strategy may limit model generalization. To address this issue, we propose SlideTTT, a slide-specific test-time training framework that enables each WSI to adapt its aggregation process before prediction. At its core, we introduce a Slide TTT Layer that constructs an inner-learning problem from patch tokens, where key-value pairs dynamically update lightweight fast weights and query tokens retrieve slide-specific representations from the adapted state. The layer combines complementary global and local adaptation: the global branch captures long-range semantic dependencies, while the local branch models short-range spatial interactions through convolutional fast weights. Moreover, the Slide TTT Layer is designed as a modular component that can be readily integrated into other Transformer-based MIL architectures. We evaluate SlideTTT on five pathology tasks using paired internal and external cohorts. SlideTTT achieves the highest average accuracy across internal, external, and overall evaluations, with more pronounced gains on external cohorts, demonstrating the effectiveness of slide-specific adaptation for improving generalization to unseen data. Code is available at https://anonymous.4open.science/r/SlideTTT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.