Patchwise 3D Context Embedding: Equipping Efficient 2D Foundation Models With 3D Perception For Brain Hemorrhage Segmentation
Abstract
Brain hemorrhage segmentation is clinically important on non-contrast computed tomography (CT), but current methods face a trade-off: task-specific 3D models capture volumetric context at high computational cost, while 2D foundation models offer efficiency and strong pretrained representations but lack cross-slice context. Here, we ask whether a 2D foundation model can be efficiently equipped with the missing 3D prior for brain hemorrhage segmentation, without relying on computationally expensive dense 3D or video-style volumetric processing. We propose Patchwise 3D Context Embedding (P3CE), an efficient approach for introducing 3D prior into 2D foundation models. Instead of densely processing the full volume, P3CE represents local 3D neighborhoods as sparse point-cloud tokens, allowing information from adjacent axial slices to be captured efficiently. We first pre-trained a brain-specific Point-MAE on CT volumes to capture volumetric anatomical context. The resulting frozen encoder generates attributed 3D patch tokens once for each case, providing reusable 3D context for subsequent slice-wise segmentation. A slice-conditioned retrieval module then selects relevant tokens for each axial slice and injects them into the 2D model through lightweight cross-attention and residual fusion. Segmentation remains slice-wise, while 3D context is encoded only once. We evaluated P3CE on Seg-CQ500 and CT-ICH. P3CE achieved the highest overlap accuracy among the compared methods (0.5686/0.4773 IoU, 0.7078/0.6343 Dice), improving Dice by 9.25 points over the SAM-Med2D baseline on Seg-CQ500 while adding only a single frozen pointencoding pass per volume, with notable improvements on small and low-contrast hemorrhages and competitive inter-slice continuity. Ablations confirmed the contributions of both the 3D point-cloud representation and point attributes. These results show that 2D foundation models can benefit from volumetric information without dense 3D computation. By separating the computation unit from the context unit, P3CE provides an efficient interface for introducing sparse 3D information into slice-wise foundation models, suggesting a potential route for extending them to other medical imaging tasks where 3D context is important, pending broader validation. This study also provides a pre-trained, brain-specific Point-MAE for head CT.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.