Same IoU, Different Decisions: Synthetic Mask Corruptions Mislead Classifiers Guided by Segmentation
Abstract
Segmentation masks guide downstream classifiers through masking, cropping, auxiliary channels, and attention. Synthetic corruptions of ground-truth masks are a convenient substitute for repeatedly training segmenters, but their use assumes that matching intersection over union (IoU) also matches downstream effects. We test this assumption with a source-crossed audit that trains and evaluates classifiers on both synthetic masks and held-out segmenter predictions, while matching the two sources to the same IoU target for each image. On frozen PlantSeg test data, the two training sources respond differently to the same change in evaluation mask, yielding a +6.28 pp difference-in-differences (95% CI over six seeds [5.31, 7.26]). A preregistered BUSI replication gives +16.50 pp ([13.93, 19.07]), whereas a CUB extension using six seeds gives a smaller +2.07 pp interaction ([1.53, 2.62]) in the opposite practical direction: synthetic evaluation fails to reveal a +1.99 pp benefit from ROI cropping with predicted masks. On PlantSeg, synthetic validation consequently selects a masking method that incurs +1.91 pp deployment regret ([0.48, 3.34]). Matching fragmentation 350× more closely does not remove the discrepancy, and training masks obtained by cross fitting only partially mitigate it. Thus matching IoU for each image does not make synthetic and predicted masks interchangeable downstream. Their fidelity must be tested with a source-crossed downstream audit rather than inferred from mask similarity
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.