CAViR: A Synthetic Video-Text Benchmark for Real-World Anomaly Retrieval
Abstract
Language-based anomaly clip retrieval searches pre-segmented short clips whose content matches a natural-language description of a normal or anomalous event. However, real-world collection, e.g., the XD-Violence dataset, faces prohibitive cost and privacy constraints, and covers only a fraction of anomalous event types. We introduce CAViR (Cross-modal Anomaly Video Retrieval), a synthetic video-text benchmark comprising 41,315 short video clips (1.36M frames) across 30 normal activity types and 68 anomalous event types. Each clip is paired with 5 captions, yielding 206,575 video-text pairs. To test whether the generated data is learnable and transferable, we further propose the FLAME (Flow-guided Language-video Alignment with Mixture-of-Experts) framework, which integrates Token-Aware Mixture-of-Experts (TAM) and a Cross-modal Flow Predictor (CFP). TAM adaptively routes visual and textual tokens to specialized experts for fine-grained event representation learning, while CFP uses auxiliary optical-flow prediction to align language with short-term visual dynamics. FLAME achieves the highest text-to-video/video-to-text Recall@1 among competitive methods: 71.6%/87.3% on CAViR, 41.0%/39.7% on UCFCrime-AR, and 33.4%/33.8% on XD-Violence, respectively. Extensive experiments verify synthetic-to-real transfer. Without relying on any real anomaly clips, pre-training on CAViR alone improves text-to-video Recall@1 on UCFCrime-AR from 22.4% to 33.8%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.