When Early Is Not Early: The Precursor Illusion in Time-Series Anomaly Detection
Abstract
Early anomaly warning requires both detectable evidence before event onset and an alarm that exploits that evidence. Yet common evaluations define pre-onset targets by temporal proximity and apply point adjustment, allowing reactive detections to receive credit as apparent predictions. We term this failure mode the precursor illusion. We introduce a model-agnostic, placebo-calibrated framework that first tests whether pre-onset intervals are distinguishable from matched normal operation, then tests whether frozen causal detector scores exploit the certified evidence, and finally isolates evaluation effects while preserving scores and thresholds. The test detects evidence in 43 of 47 independently annotated transitions in the 3W dataset, yet across five benchmarks it identifies pre-onset deviations in only 70 of 500 evaluable events, revealing sparse and dataset-dependent evidence within the specified descriptor family. Mean detector scores achieve above-chance discrimination between precursor and matched normal windows in none of the 35 model–dataset combinations after multiple-testing correction, and only the near-onset score peaks of LSTM-VAE separate SMD precursor windows beyond a selection control, although all seven evaluated models distinguish the 3W transitions from normal operation. Point adjustment raises mean pre-onset Shifted-warning F1 from 0.036 to 0.177 without changing any alarm time, and input-blind alarms at matched rates obtain higher adjusted F1 than every detector. These results show that benchmark warnability must be established independently before early-warning models are ranked.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.