acceptodds
Under review as a conference paper at ICLR 2027

Every Scale Has Its Geometry: Rethinking Positional Encoding for Multi-Resolution Multiple Instance Learning

Abstract

Multiple Instance Learning (MIL) enables weakly supervised analysis of Whole Slide Images (WSIs) without costly patch-level annotations, but its patch-based formulation inherently disrupts the spatial organization of tissue structures. Positional encodings (PEs) have therefore been introduced to restore spatial information lost during patchification. More recently, Multi-Resolution MIL (MRMIL) has leveraged complementary information across magnifications by incorporating patches at multiple resolutions, better reflecting the multi-scale visual assessment performed in clinical diagnosis. However, existing PEs are not well suited to the geometry of multi-resolution WSIs (MR-WSIs): they either model spatial relationships independently within each resolution or inadequately account for the heterogeneous spatial scales induced by different magnifications. In particular, topology-aware PEs typically treat cross-resolution spatial relationships uniformly, even though the spatial domain represented by patches expands quadratically with increasing magnification. To address this limitation, we propose Hyperbolic Hierarchical Positional Encoding (H2PE), a positional representation that jointly captures spatial topology and resolution hierarchy within the hyperbolic upper half-space. By leveraging the exponential capacity of hyperbolic geometry, H2PE naturally represents the hierarchical, scale-dependent structure of multi-resolution patches while preserving their spatial relationships across magnifications. H2PE is model-agnostic and can be incorporated into arbitrary MRMIL aggregation frameworks without modifying their underlying aggregation architectures. Extensive experiments on three public benchmarks demonstrate consistent improvements over existing positional encodings across multiple MRMIL frameworks. Further geometric analyses and visualizations validate that H2PE effectively captures the hierarchical organization and cross-resolution spatial structure of MR-WSIs, while providing interpretable positional representations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.