MaskMed: Decoupled Mask and Class Prediction for Medical Image Segmentation
Abstract
Medical image segmentation methods typically use point-wise convolutional heads that map output channels to predefined semantic classes. While effective, this formulation imposes a rigid channel-to-class correspondence that limits flexible spatial-semantic reasoning and feature reuse across classes. We propose MaskMed, a volumetric segmentation framework that decouples mask generation from class prediction. Instead of directly predicting multi-class logits, MaskMed produces class-agnostic binary masks together with their semantic labels using a shared set of object queries. This formulation removes fixed class-channel bindings and enables more adaptive semantic reasoning over anatomical structures. To improve multi-scale feature integration, we introduce a Full-Scale Deformable Transformer that enables deformable cross-scale interaction between encoder and decoder features across the full feature hierarchy while preserving spatial alignment. We further design a masked multi-scale segmentation head with shared query propagation and finest-scale bipartite supervision for consistent coarse-to-fine refinement across decoder stages. Extensive experiments on AMOS22, BTCV, ACDC, and BraTS23 demonstrate state-of-the-art performance. MaskMed improves Dice scores over nnU-Net by +2.0 and +1.3 points on AMOS22 Task 1 and Task 2, respectively, as well as by +1.7 on ACDC, +0.9 on BTCV, and +1.0 on BraTS23.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.