Fourier Principal Projected Transformer for Efficient and Accurate Medical Image Segmentation
Abstract
Accurate medical image segmentation requires reliable modeling of global anatomical structure in high-resolution medical images. However, existing Transformer-based methods lack explicit global structural representation and suffer from quadratic attention complexity, limiting their effectiveness and efficiency in dense prediction settings. To address these limitations, we propose the Fourier Principal Projected Transformer (FPP-Former). Specifically, FPP-Former leverages the intrinsic global correlations of spatial tokens to derive principal component representations that characterize the major variations in global anatomical appearance, enabling explicit modeling of global anatomical organization. For efficient principal component modeling, we reformulate this process in the Fourier domain, where a low-dimensional principal subspace is constructed from dominant frequency components, avoiding explicit eigen-decomposition on high-dimensional feature representations. The resulting Fourier principal components are then fed into the attention model to guide key–value interactions, where keys and values are projected into a shared principal space for structure-guided attention computation and value aggregation, thereby focusing attention on spatially continuous and structurally coherent anatomical regions. This formulation results in linear-complexity attention computation, improving efficiency for high-resolution segmentation. Extensive experiments on BraTS 2017, Synapse, and ACDC show that FPP-Former achieves average Dice scores of 83.2%, 84.84%, and 91.94%, with relative improvements of 2.34%, 0.11%, and 0.12%, respectively, while consistently improving training and inference efficiency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.