Predict-Then-Adapt: Cross-Plane Joint-Embedding for Test-Time Medical Segmentation
Abstract
Test-time medical image segmentation adapts source-trained models to unlabeled target volumes under distribution shifts. Existing methods mainly derive adaptation signals from model predictions or individual slices, leaving the intrinsic three-dimensional structure of test volumes underexploited. In particular, orthogonal views observe the same anatomy with exact voxel-wise correspondence, providing natural within-volume supervision. However, a representation that is predictable across views may not remain compatible with the source-trained segmentation task. We propose CP-JEPA, a Cross-Plane Joint-Embedding Predictive Adaptation framework following a predict-then-adapt principle. CP-JEPA models cross-plane prediction and task-space translation as two complementary stages. It first predicts masked target-view representations from geometrically corresponding orthogonal views, constructing a JEPA-style objective within each test volume. It then compares the cross-plane prediction with a capacity-matched axial control to isolate geometry-specific counterfactual evidence and translates this evidence into bounded, source-anchored segmentation corrections. CP-JEPA supports both single- and continual-volume adaptation by resetting volume-specific predictive states while selectively retaining reliable task-space evidence across cases. We evaluate CP-JEPA on multi-organ MRI segmentation under single- and continual-volume deployment, with and without labeled target-development data used only for configuration. Across DINOv3- and U-Net-based backbones, CP-JEPA achieves the strongest macro-average performance in three of four evaluation settings and ranks second in the remaining setting, outperforming the strongest TTA baselines by approximately Dice point on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.