From Detection to Translation: Rethinking Evidence Sampling for Video Anomaly Understanding
Abstract
Video Anomaly Understanding (VAU) aims to identify abnormal events in videos and explain their visual evidence, event process, and semantic causes in natural language. Existing VAU pipelines commonly rely on uniform sampling or anomaly-score-based filtering before Multimodal Large Language Model (MLLM) reasoning, where subtle but critical abnormal frames may be removed before the reasoning stage. To address this challenge, we propose a translation-based sampling framework: instead of making coarse anomaly decisions during sampling, our method translates all video clip units into compact visual-semantic representations and performs text-driven redundancy merging to retain representative clips with their corresponding text priors. We further integrate scene context into motion-aware translation and accelerate evidence construction via dynamic token pruning to improve both accuracy and efficiency.Experiments on HIVAU-70k show that, compared withHolmes -VAU baseline, our method improves average BLEU and ROUGE performance by 13.4% and 20.4%. The codes will be public after acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.