When High ROC-AUC Fails: Score Collapse and the Gradient Exposure Deficit in Anomaly Detection
Abstract
In industrial anomaly detection, a detector can rank short anomalous sensor segments accurately yet produce unreliable alarms for complete operating runs when conditions change. We study this on a differential-pressure pipeline leak-detection dataset and three public time-series benchmarks. The failure mechanism is score collapse: under binary cross-entropy (BCE) training with class imbalance, normal-class scores concentrate near zero, making the alarm cutoff unreliable while the ROC-AUC (area under the receiver operating characteristic curve) remains unchanged. Score collapse produces two distinct threshold failures—cutoffs that saturate to produce false alarms on every run, and cutoffs that miss nearly all leak-containing runs—both of which pass a standard high-AUC evaluation checkpoint. We propose score-histogram entropy (, nats; entropy computed from natural logarithms) as a complement to ROC-AUC, and evaluate four failure modes requiring distinct interventions: score dispersion, probability calibration, sensitivity to head-loss modifications, and alarm-cutoff performance under operating-condition shift. Joint entropy regularisation restores score dispersion ( rises from to nats on the pipeline dataset); post-hoc Platt scaling reduces expected calibration error from to –; and density-ratio weighted conformal prediction improves run-level recall from to at zero false alarms in a stress test under a shifted operating condition. Score collapse is reproduced across three public benchmarks and three encoder architectures; gradient exposure—not backward-call placement per se—is the operative factor in collapse prevention. Score dispersion should be evaluated alongside ROC-AUC before a detector is used in a thresholded deployment setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.