From Prior to Precision: Unifying Global Structural Priors and Local Generative Refinement for Medical Image Segmentation
Abstract
Medical image segmentation is acutely bottlenecked by the scarcity of pixel-level annotations, especially in the low-data regime. Vision foundation models(VFMs) such as DINO provide rich semantic and structural priors that can alleviate this dependency, yet cross-domain discrepancies erode their direct transfer to the fine-grained representations required by medical segmentation. In this work, we propose **P**rior-t**o**-**P**recision **Net**work (PoPNet), a unified framework that bridges generic foundation-model priors with task-specific segmentation precision through global structural regularization and local generative refinement. To capture global structures, PoPNet uses a spectral clustering regularization to anchor the task-specific latent manifold toward structural priors derived from VFM cross-layer attention affinities, encouraging the learned representation to preserve dominant structural patterns while reducing reliance on domain-specific feature details. To further improve local precision, PoPNet incorporates target-aware region-of-interest masked generation, which performs conditional reconstruction within target regions, compelling the network to capture fine-grained textures and subtle boundary details critical for accurate delineation. Extensive experiments across five supervised benchmarks and semi-supervised segmentation settings demonstrate that PoPNet achieves state-of-the-art performance with strong generalization across datasets and foundation-model priors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.