JELI: Joint Modeling of Latent and Input-Space Evidence for Time Series Anomaly Detection
Abstract
Time-series anomaly detection is a crucial task, commonly formulated as learning characteristic temporal patterns from normal observations and identifying abnormal deviations. Prior work has broadly explored time-series anomaly detection in two spaces, with input-space approaches detecting anomalies through discrepancies between inputs and their reconstructions, while latent-space approaches identify deviations from normal patterns in learned representations. However, the two approaches have largely evolved as separate research directions, even though they can serve complementary roles in anomaly detection. In this paper, we propose **JELI** (**J**oint **E**vidence from **L**atent and **I**nput Spaces), a unified framework that integrates contextual anomaly evidence from both spaces for time-series anomaly detection. Specifically, we first train a JEPA-based backbone to learn contextual representations from surrounding input context and predict target representations at arbitrary temporal positions, forming a normal representation space. Second, we finetune the pretrained backbone for contextual reconstruction with an attention-masking bottleneck that prevents each target from attending to its own representation while preserving alignment with the pretraining objective. Lastly, to bridge the two sources of anomaly evidence, we design a joint-distribution-based scoring method that explicitly accounts for their dependence. Through extensive experiments, we demonstrate that our framework substantially outperforms prior state-of-the-art baselines on the time-series anomaly detection benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.