Open-Ontology Gap in Free-Text Medical Segmentation
Abstract
Free-text medical segmentation lets users request clinical targets in ordinary language, but a plausible mask may represent the wrong concept. Prior work reports category errors and limited generalization, but lacks a controlled test of whether a model selects the requested target when a related alternative is present. We introduce Clinical Contrast Sets, which holds the image fixed, changes the requested concept, and measures whether the mask shifts toward the intended region and away from its competitor. On 40 MS3SEG development patients, Medical-SAM3 and VoxTell preferentially select multiple sclerosis (MS) lesions when asked for the dataset’s normal/non-MS white-matter hyperintensities (nWMH). A protocol frozen before evaluating the 20 official-test patients confirms this pattern: mean target-minus-confounder coverage is −0.308 and −0.363, with 0/20 correct canonical prompt pairs for each model. Positive descriptions and volume controls preserve the failure, local size matching confirms it for VoxTell, and published supervised results establish that nWMH is learnable. The failure is not a missing label: VoxTell’s published training inventory lists both MS lesions and white matter hyperintensities, and the model switches correctly between documented kidney and tumor concepts, yet it resolves nWMH prompts, and even the generic white-matter-hyperintensity prompt, to MS lesions. Incorrect nWMH masks retain a mean foreground probability of 0.962, and a prompt-overlap warning that flags all 20 official VoxTell failures performs near chance within BraTS. This matters because the failure is not a missing answer but a wrong one: a clinician requesting a benign, non-MS finding instead receives a confident, well-formed mask of MS pathology, with nothing in the model’s own output to flag the substitution. We call this mismatch between language flexibility and reliable visual target selection the open-ontology gap, and offer Clinical Contrast Sets as a controlled protocol for auditing it: with at least two annotated competitors, two prompts, and a superclass reference, developers can identify which clinical distinctions a model does not reliably select and verify whether an adaptation restores correct switching rather than only raising overlap.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.