Where Does Detection Fail? A Residual-Evidence Diagnostic Workflow for Industrial Control Systems
Abstract
Residual-based anomaly detectors for industrial control systems (ICS) convert model residuals into anomaly scores and alarms, but final detection metrics give little guidance on which component to examine when detection performance is poor. We present a retrospective diagnostic workflow built around covariance-normalized residual energy (CNRE), a measurement based on the squared Mahalanobis distance. The workflow asks whether the normal reference still describes the evaluation recording, whether attacks separate from normal operation under a stated reference, and how scoring and alarm rules use fixed residuals. Across fifteen detector-dataset pairs (five detectors on SWaT, WADI, and HAI), reference updates improve separation in some pairs but degrade it in others, and a retrospective policy that updates only where a label-free screen recommends it raises mean AUC from 0.751 (never updating) and 0.759 (always updating) to 0.809; on the twelve non-GDN pairs, which did not inform its boundaries, it gives 0.805 versus 0.759 and 0.747. Rescoring CNRE computed with a fixed reference reproduces most of the AUC gain of a reference update on WADI but not on SWaT. A detailed GDN case study connects the diagnostics: normal-only adaptation reduces normal prediction error by 27.3%, but the direction of its effect on attack separation depends on the reference protocol. With residuals fixed, stronger score ranking does not ensure that normal-calibrated alarms preserve their nominal false-positive rate. These results support choosing intervention stages and interpreting their outcomes through separate measurements of reference fit, attack evidence, and alarm calibration.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.