Cross-Scale Latent Predictive Pretraining for Whole-Slide Image
Abstract
Whole-slide images (WSIs) contain diagnostic information across magnifications, from tissue architecture to cellular morphology. Multi-scale MIL methods often encode patches from different magnifications independently and combine their representations post hoc, so the coarse representation may fail to preserve cues predictive of fine-scale morphology. Using both scales also requires fine patch encoding at inference. Cross-scale self-supervised learning (SSL) introduces interaction during pretraining, but direct representation matching or fine-patch reconstruction does not account for cross-scale asymmetry because fine views contain details unavailable from coarse inputs. We introduce **Patch-JEPA**, a cross-scale latent predictive pretraining framework. A frozen fine-scale target encoder produces an ordered fine-scale latent field from aligned fine patches. Separate prediction branches map coarse representations to this field and an aggregated coarse-scale target, training the coarse-scale context encoder without direct cross-scale matching or fine-patch reconstruction. Fine-scale representations are used only as pretraining targets, and the retained context encoder uses only coarse patches for coarse-only inference. Patch-JEPA is competitive with standard SSL, cross-scale SSL, and post-hoc multi-scale WSI pipelines on four pathology benchmarks. On RCC, ridge probing yields \(R^2 = 0.450\) for aligned fine-scale latent-field prediction, compared with \(0.231\) for Cross-MAE. Model-side inference cost remains close to a single-scale baseline under a fixed-bag timing protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.