Pretrained Systems for Low-Data Myocardial Segmentation: Direct Versus LV-Supervised Fine-Tuning
Abstract
Choosing a pretrained system is important when myocardial annotations are limited. We compare Pixio-L with Mask2Former and EchoFM with UNETR for myocardial (MYO) segmentation on CAMUS apical four-chamber (A4C) echocardiography, evaluated against official provider-supplied reference annotations. Pixio-L is a released student distilled from the web-scale Pixio-5B teacher; EchoFM is pretrained on echocardiographic videos. Each system starts from its released encoder weights and is trained either directly on MYO or first on EchoNet-Dynamic A4C left ventricular (LV) cavity segmentation and then on MYO. Under full fine-tuning, across target-training patient fractions of 5%, 10%, 25%, 50%, and 100% and three training seeds, Pixio-L achieves higher mean MYO Dice under both routes. At 5%, the cross-system differences favoring Pixio-L are 13.38 percentage points for Direct fine-tuning and 6.52 points after LV supervision. The LV stage improves mean Dice in both systems at every fraction, with the largest gains at 5%: 2.29 points for Pixio-L and 9.15 for EchoFM. An additional 60-run frozen-MYO-encoder ablation shows that the effect of encoder updates depends on the system and initialization route. These exploratory results on one fixed patient split identify Pixio-L with Mask2Former as a strong low-data baseline and support intermediate LV supervision for both systems. Differences in architecture, input, and training recipe prevent attributing the ranking to pretraining domain or distillation alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.