acceptodds
Under review as a conference paper at ICLR 2027

MemVAU: Agentic Visual Evidence Memory for Video Anomaly Understanding

Abstract

Video Anomaly Understanding (VAU) explains abnormal events along four dimensions: What, Who, Where, and How. Existing single-pass vision-language models (VLMs) do not explicitly check visual support for each dimension, leaving difficult fields such as Who and How insufficiently supported by localized visual evidence. We introduce MemVAU, an agentic visual evidence memory framework that recasts VAU as a closed loop. An Observer generates candidate evidence, a frozen Verifier scores it along the four dimensions, and a rule-based Mem-Optimizer retains complementary evidence in a bounded pool. If any dimension remains below its operational threshold and the round budget is not exhausted, the structured pool snapshot guides the next evidence-acquisition round. The Reporter synthesizes the final report when all dimensions meet their thresholds or the budget is exhausted. The Observer is trained with GRPO under a listwise reward for anomaly-type accuracy and balanced coverage. On HIVAU-70k, MemVAU achieves the best score on 10 of 12 metrics, with its largest gains at the Event and Video granularities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.