acceptodds
Under review as a conference paper at ICLR 2027

Knowing When Not to Segment: Label-Free Adaptation in Free-Hand Ultrasound

Abstract

Self-training and other label-free adaptation methods for segmentation usually assume that every image contains the target structure, but in videos and sweeps the structure is often out of view. Wireless handheld probes bring ultrasound out of the hospital, but many of their users are not medical professionals and cannot judge whether a model's prediction is correct. Segmentation models trained on public data from cart-based scanners segment structures in handheld images well when they are in view, but they do not know when not to segment. We study how to teach such a model, without target labels, to predict nothing when a structure is not in the frame. We built a dataset of free-hand thyroid sweeps with a wireless handheld probe, which we will release. It contains 20,014 frames from 29 adults, 232 of which are annotated for thyroid, carotid artery and trachea. Of these, 47% show no identifiable target structure, and two clinicians annotated the 58 held-out frames of its hidden test set. Our experiments reveal two problems. First, current evaluation measures only how accurately a structure is drawn and ignores frames without it, so it cannot show whether a model chooses not to draw when the target structure is absent. Second, because existing models are trained on datasets in which nearly every image contains the target structure, they also draw it on frames where it is absent. We therefore propose an evaluation that also scores predicting nothing. Frames with a structure are scored by how accurately it is drawn and frames without it by whether nothing is drawn, with equal weight, so predicting nothing scores 0.50. We also add anatomical constraints to self-training. Predicted structures are checked against rules on their position, shape, echogenicity and relations to the other structures, and implausible predictions are not used for training. The model thereby learns to predict nothing. On the evaluation and held-out frames, false thyroid detections fall from 99% and 90% to 40% and 29%, and the score reaches 0.59 and 0.61, significantly above predicting nothing and above the seven other adaptation methods that start from the same source models, while mean segmentation accuracy on frames with the structure does not drop.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.