acceptodds
Under review as a conference paper at ICLR 2027

Predictive Transitions for Video Anomaly Detection

Abstract

We derive a direct anomaly signal from a representation learned to condition future-frame generation, using a state–transition framework for normal-only video anomaly detection. Current appearance and ordered frame and feature changes jointly condition a diffusion Transformer. After generative training, we freeze the encoders and fit a small head to predict the next local frame change from transition features. Matched representation probes on Ped2, Avenue, and ShanghaiTech show that learned transition features outperform raw frame-change history for anomaly detection. Interventions reveal dependence on transition content and temporal order, with clearer temporal-order effects on Ped2 and Avenue. Together, the probes and conditioning ablations show that transition features can support anomaly detection even when hybrid conditioning does not improve frame-error AUC. The change score can be computed from the learned encoders and prediction head without iterative frame generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.