PV-WSI: Rethinking Pseudo-Video Learning for Whole Slide Image Segmentation via Spatial Discontinuity Modeling
Abstract
Whole-slide image (WSI) segmentation for computational pathology is challenged by gigapixel scale, extreme morphological heterogeneity, and fragmented cross-patch context. Recent pseudo-video formulations attempt to treat WSI patches as sequential frames, but inherit a critical flaw: they impose temporal smoothness priors from natural video onto inherently discontinuous pathological landscapes. This mismatch leads to erroneous context propagation and degraded segmentation, especially across structurally dissimilar regions. We introduce PV-WSI, a pseudo-video paradigm that explicitly models both spatial continuity and discontinuity. Unlike prior work that passively adapts video architectures, PV-WSI reconstructs pseudo-temporal structure. Specifically, we propose Adaptive WSI Pseudo-Video Construction, which identifies pathology-informative anchor patches and dynamically assembles morphology-consistent sequences, eliminating traversal bias and enforcing structural alignment. Furthermore, we develop Distance-Constrained Gate State Propagation, a principled distance-aware propagation mechanism that selectively transmits reliable contextual states while suppressing spurious long-range dependencies across weakly correlated regions. Extensive experiments show that PV-WSI consistently outperforms state-of-the-art methods across magnification levels, improving both tumor segmentation and multi-class tissue parsing. These results demonstrate that explicitly modeling spatial discontinuity is essential for pseudo-video learning and support PV-WSI as a strong framework for large-scale WSI understanding. Code will be available at GithubGithub.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.