Same Dice, Different Mistakes: Label Coarsening Relocates Segmentation Errors
Abstract
Segmentation metrics such as Dice and IoU report how much a model gets wrong, but not what it gets wrong: they discard the identity of the pixels that receive false positives. That is harmless only if false positives are interchangeable; they are not. A multiple sclerosis (MS) lesion detector that flags a benign white-matter hyperintensity makes a different, clinically consequential mistake than one that flags ordinary tissue; a car detector that fires on a truck differs from one that fires on road. Training pipelines erase this distinction by collapsing every non-target structure into one background class. Prior work has shown that keeping such structures separate can raise Dice. What has not been shown is that coarsening changes the composition of a model's errors even when their quantity and Dice are held fixed: when coarse- and fine-labeled models make the same number of mistakes, do they make the same kinds of mistakes? We show they do not, and name the phenomenon semantic error relocation (SER). We introduce Semantic False-Positive Concentration (SFPC), the fraction of target false positives that fall on a pre-specified hidden negative subclass, and a protocol that compares models only after matching their total false-positive burden on validation data, then under an exact equal-error-budget audit in which any SFPC difference is a relocation of fixed error mass. On MS3SEG, in a pre-registered confirmation across three 3D architectures and three seeds, binary supervision places 20.96% of its false positives on normal white-matter hyperintensities versus 13.81% under four-class supervision (+7.15 percentage points, 95% CI 6.04–8.31, 39/40 patients), while total false positives are indistinguishable (1280 vs. 1283 voxels) and binary Dice is slightly higher (0.626 vs. 0.617). This is about one additional benign region falsely flagged per patient, a 266× enrichment over the subclass's area. An independent Cityscapes replication reproduces the effect: +7.24 points more car false positives land on trucks and buses (95% CI 5.09–9.59) at statistically identical Dice (0.865 vs. 0.867). Controlled follow-ups rule out prediction-head size as the cause and trace the shift to the training objective the taxonomy induces: reweighting the hidden class alone reproduces part of the effect, with any further contribution from explicit class identity left open. The practical conclusion is direct: whenever fine labels exist at evaluation time, even if deployment needs only a binary mask, reporting SFPC alongside overlap metrics exposes a difference in what errors mean that Dice cannot see.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.