What Video Anomaly Probes Measure
Abstract
An above-chance probe score on frozen video features is often read as evidence that the representation captures anomaly. We argue that this reading needs an evaluation contract, and we audit the supervised probe with four low-dimensional controls—scene complexity, video identity, position in video, and detected-agent counts—under scene-grouped folds, paired rows and video-clustered uncertainty, on a frozen V-JEPA 2 encoder over UCF-Crime and NWPU-Campus. On UCF-Crime, a linear probe given only the first encoder window of each clip predicts whether the clip will contain an anomaly at video-level AUROC 0.899; this is a video-level diagnostic, not evidence of temporal anticipation. On NWPU-Campus, a nine-dimensional count vector is a strong comparator under the diagnostic: with each arm's better of two heads, no tested representation arm is superior after paired Holm correction, and the count arm is the highest selected arm in every one of 25 stored seeds. That ordering is conditional on the readout: under a fixed linear head the count arm sits at chance and no arm is distinguishable from it, so we report paired tests within each head family rather than a single ranking. The seed sweep jointly varies fold assignment and optimisation and moves arms by up to 0.186 AUROC, widest for the complexity control. Under the benchmark's own normal-only protocol the same counts reach at most 60.6 micro-AUC against a published 76.9, whereas the frozen state reaches 59.1 micro-AUC and an exploratory 67.2 macro-AUC in the prespecified estimator cell: a control can dominate a supervised diagnostic without being a competitive detector. Two analytical results—that a prediction-error residual is a dispersion functional, and that a self-normalised pre-hoc hazard measures the model's own tail mass—motivate the complexity control and explain why no unsupervised anticipation arm appears. We release the protocol and every artefact, and limit the conclusion to measurement design: an anomaly score should state its target, information cutoff, controls, grouping, readout selection, aggregation and uncertainty before it is read as evidence about a representation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.