Detecting Time-Series Anomalies with Score Models Alone
Abstract
Diffusion and score-based generative models expose the gradient of the log-density, , not the density itself. Can such score-only models detect anomalies in time series, and what exactly can they detect? First, the capability: a directly parameterized score network, scored by the window-averaged Hyvärinen statistic , reaches VUS-PR on the TSB-AD univariate benchmark, with no statistically significant difference from the strongest baseline PaAno (which stays ahead under a buffer-free control), at under a third of its measured runtime; on the multivariate suite it is mid-field. The divergence term is load-bearing and significant: over the same network's score norm (). Second, the limits: we prove the population detection gap equals one half the Fisher divergence plus a signed Fisher-information correction, naming a blind spot of the upper-tail test that more data cannot remove: anomalies sharper than normal data invert the statistic whenever , and we observe an inversion consistent with this class on a real sensor flat-line. Third, the mechanism: window aggregation estimates the statistic's population expectation (i.i.d. window model), and footprint-scale aggregation () performs near the best tested scale in our synthetic length-group averages. Fourth, an equal-treatment audit: the same aggregation lifts five saved-score baselines by – VUS-PR and reorders the leaderboard (the lift surviving a buffer-free control), so part of the apparent ranking is output post-processing. In our matched energy-network comparison on the univariate suite, the two statistics show no statistically significant difference; the score-only statistic is useful when no energy readout is available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.