UniAR: Benchmarking Cross-Modal Evidence-Grounded Anomaly Reasoning for MLLMs
Abstract
Anomaly reasoning is a core perceptual and cognitive ability in humans, who can identify anomalies across distinct modalities and provide corresponding evidence. Nevertheless, there remains a lack of a unified benchmark to rigorously and credibly evaluate such ability in Multimodal Large Language Models (MLLMs) on comprehensive tasks spanning video, image, and text modalities. To this end, we present UniAR, the first cross-modal anomaly reasoning benchmark with full-spectrum tasks and credible evaluation. Built from publicly available source data, UniAR provides over 40,199 high-quality question-answer pairs organized into four shared challenges: dual anomaly discrimination, context understanding, fine-grained anomaly classification, and holistic anomaly analysis, covering coarse-to-fine anomaly reasoning across video, image, and text modalities. To improve credibility, we go beyond standard multiple-choice questions (MCQs) by assessing the evidence MLLMs provide for their answers, measuring whether they genuinely comprehend the anomalous phenomena. We benchmark over 20 latest models, and find that though proficient in semantic understanding and commonsense reasoning, they show clear modality bias in anomaly discrimination and lag substantially behind human performance across all tasks, indicating considerable room for improvement. We hope UniAR can facilitate research on human-like anomaly reasoning for embodied agents and intelligent assistants.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.