AnomMerge: Spatially Bounded Token Merging for Video Anomaly Understanding
Abstract
Video anomaly understanding (VAU) with multimodal large language models (MLLMs) aims to recognize, temporally localize, and explain abnormal events. Under limited visual token budgets, existing VAU systems often trade temporal coverage for spatial detail, potentially removing evidence needed for precise event boundaries. Separately, we observe that current VAU models can produce overly broad and low-dispersion anomaly intervals, while standard overlap metrics may not fully expose this degeneration. To address these challenges, we propose AnomMerge, a learnable token compression framework for VAU that reduces spatial redundancy while preserving observed temporal units and their positional correspondence. We further develop adaptive spatial merging with unit-specific similarity thresholds and spatial and feature consistency constraints to preserve fine-grained temporal evidence for anomaly localization. To better evaluate localization behavior, we complement standard overlap metrics with coverage-aware temporal IoU (CA-tIoU) and interval dispersion contraction (IDC), which measure excess temporal coverage and contraction of predicted interval dispersion, respectively. On Vad-Reasoning, AnomMerge reduces IDC by 44.3% relative to VAU-R1, while improving video-level anomaly recognition and retaining only 36.59% of the original visual tokens. The code will be publicly released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.