How Much of Anomaly Segmentation Is Background? A Published 97% Pixel AUROC Can Certify Chance Inside the Object
Abstract
A published 97% pixel AUROC on liver CT certifies chance inside the organ. Anomaly segmentation, medical and industrial alike, is scored almost universally by that number over the whole image, and we show what it measures. On liver CT a Gaussian blur of the input, with no model, no training and no support set, scores 94.8% and a binary foreground mask 96.9%, within three points of published detectors; inside the organ they score 14.0% and chance. The reason is an identity: whole-image AUROC equals w*A_B + (1-w)*A_F, weighted by the fraction w of normal pixels that are background, so a reported score certifies only A_F >= (A-w)/(1-w); on liver CT w=0.94, and the claim is about w: on breast ultrasound (w=0.15) the same 97% certifies A_F >= 96.5%. Class imbalance is not the mechanism, and balancing the classes leaves the number unchanged; the partition is a reporting choice that leaves A untouched, and across 6 partition rules the liver-CT guarantee moves by 2.6 points. The clearest evidence is a single detector: masking the background out of a nearest-neighbour memory takes its whole-image score on liver CT from 97.1 to 4.5 while its foreground score barely moves, 64.7 to 63.4 — same weights, same images, only the negative class changed. Measuring every term on 53,260 images across 8 dataset families spanning medicine, industry, remote sensing and driving scenes, the inflation follows w wherever a predictor can separate the background and vanishes where it cannot. Re-running training-free detectors on the six medical benchmarks under both metrics, the ranking they induce changes wherever the frame is empty (liver CT: Kendall tau=0.12, 40 of 91 pairs flip) and holds where it is full (breast ultrasound: tau=0.98). We recommend reporting A_F beside A with the partition stated, no new metric but the same statistic on a stated set of normals, release the audit as a one-file tool, and report a training-free reference baseline both ways.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.