MemVAU: Agentic Visual Evidence Memory for Video Anomaly Understanding
Abstract
Video Anomaly Understanding (VAU) explains abnormal events along four dimensions: What, Who, Where, and How. Existing single-pass vision-language models (VLMs) do not explicitly check visual support for each dimension, leaving difficult fields such as Who and How insufficiently supported by localized visual evidence. We introduce MemVAU, an agentic visual evidence memory framework that recasts VAU as a closed loop. An Observer generates candidate evidence, a frozen Verifier scores it along the four dimensions, and a rule-based Mem-Optimizer retains complementary evidence in a bounded pool. If any dimension remains below its operational threshold and the round budget is not exhausted, the structured pool snapshot guides the next evidence-acquisition round. The Reporter synthesizes the final report when all dimensions meet their thresholds or the budget is exhausted. The Observer is trained with GRPO under a listwise reward for anomaly-type accuracy and balanced coverage. On HIVAU-70k, MemVAU achieves the best score on 10 of 12 metrics, with its largest gains at the Event and Video granularities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.