acceptodds
Under review as a conference paper at ICLR 2027

ReMem-VAD: Connecting Event Representations and Memory for Anomaly Detection

Abstract

Video anomaly detection (VAD) is challenging because visually similar actions can correspond to either normal or abnormal events. Recent vision-language models (VLMs) provide broad semantic knowledge, but recognizing an action alone is insufficient when its meaning depends on the participant’s response, the event’s progression, and the evidence supporting interpretations. To address this issue, we propose to transform resuable experience to event representations and structural memory, enhancing VLMs' capabilities with memory-ground reasoning. Specifically, Anomaly Agent is proposed, a memory-augmented VAD framework is built to combine large-model understanding with lightweight anomaly scoring. For each video window, a frozen VLM represents the actions, responses, event phase, and supporting evidence. This event representation specifies what must be explained and guides the verification of competing explanations against the same frames. Under this guidance, candidate patterns are retrieved from memory and matched to the current visual evidence. A lightweight trainable scorer then learns how to combine evidence, explanation judgments, and memory pattern into an anomaly score. Counterexample-guided refinement updates memory and representation during development, and freezes them at test time. Experiments on XDViolence, UCF-Crime and UBnormal demonstrate state-of-the-art performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.