acceptodds
Under review as a conference paper at ICLR 2027

ARIA: Adaptive Reasoning for Video Anomaly Detection with Vision-Language Model

Abstract

Vision-language models (VLMs) have demonstrated strong reasoning capabilities for video anomaly detection (VAD), enabling both anomaly prediction and human-understandable explanation. Existing VLM-based methods typically process videos segment by segment, allocating similar computation to each segment despite their varying needs for further examination. While some segments can be reliably handled with sparse visual observations, others may benefit from additional evidence, either by examining denser visual observations or by incorporating a complementary interpretation from another VLM. Uniformly applying such additional computation to all segments, however, can incur substantial inference cost. To address this issue, we propose ARIA (Adaptive ReasonIng for Video Anomaly Detection), an adaptive reasoning framework that selectively allocates additional computation to video segments when needed. Given a segment, ARIA first assesses the reliability of its initial lightweight VLM response using response-level signals together with temporal and semantic context. For segments warranting further examination, ARIA then determines whether to acquire richer visual evidence through denser frame sampling or introduce complementary textual evidence from another VLM. Experiments on UCF-Crime and XD-Violence demonstrate that ARIA improves anomaly detection and explanation quality while reducing inference cost, yielding a favorable performance-computation trade-off.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.