Aggregation Scale Shapes Anomaly-Duration Bias: Bounds, Readout Effects, and a Nonparametric Detector
Abstract
Anomaly duration can change which time-series anomaly detectors perform best, yet aggregate benchmark scores obscure this variation. Across 40 published detectors on TSB-AD, relative performance varies with duration even after controlling for file fixed effects, and all three pretrained forecaster families favour short anomalies. Under stationary normal increments, a shift-invariant statistic reading w consecutive increments has expected coverage at most min1, 2w/L + α on a flat level shift of length L at false-alarm rate α: detecting an anomaly's presence is distinct from covering its interior. Short-anomaly sweeps reveal empirical dilution under the tested aggregations. Controlled interventions on two frozen pretrained forecasters separate conditioning context, forecast horizon, and score aggregation. Over the tested grid, widening aggregation from 1 to 64 samples at a fixed one-step horizon extends the post-peak half-maximum duration by roughly an order of magnitude in matched configurations with nonzero peak coverage; context has smaller effects. Increasing horizon yields a similar extension. The combined effect is sub-additive on the log-duration scale, and wider readouts reduce peak coverage. On TSB-AD-U, the same aggregation substantially reduces the relative short-anomaly bias of both frozen forecasters. These findings motivate a nonparametric multiscale detector that combines calibrated short- and wide-scale evidence by a maximum. Without gradient training, it achieves competitive offline performance against tuned deep models on TSB-AD-U and TSB-AD-M under the reported development protocol. The results identify aggregation and readout as practical controls on duration bias and support duration-aware evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.