acceptodds
Under review as a conference paper at ICLR 2027

VAD-TTRL: Propose, Verify, Reinforce for Label-Free Test-Time Video Anomaly Detection and Understanding

Abstract

Multimodal Large Language Models (MLLMs) have shown strong capabilities in video understanding, yet video anomalies are often scene- and deployment-dependent, limiting the reliability of fixed pretrained models across environments. Conventional adaptation typically requires task-specific annotations and expensive post-training. Test-time reinforcement learning offers an appealing alternative by adapting models directly on unlabeled test videos using self-generated supervision. However, existing approaches commonly treat self-consensus among sampled predictions as supervision, which can reinforce systematic errors. We introduce VAD-TTRL, a label-free test-time reinforcement learning framework based on a propose–verify–reinforce paradigm. Multiple stochastic reasoning trajectories first form a structured consensus hypothesis, which is then verified using temporal and spatial consistency together with counterfactual visual evidence. Only sufficiently supported hypotheses contribute to reliability-weighted reinforcement signals, while uncertain predictions are excluded from adaptation. These rewards are optimized with Group Relative Policy Optimization (GRPO), enabling adaptation without ground-truth anomaly labels or task-specific annotations. Experiments across multiple MLLM backbones show that VAD-TTRL improves anomaly detection and temporal localization while largely preserving reasoning quality. On Qwen2.5-VL-7B, it improves accuracy from to and mIoU from to ; when applied on top of sota Vad-R1, it further improves accuracy from to and mIoU from to , demonstrating the value of evidence-verified consensus for label-free test-time video anomaly understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.