EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
Abstract
Multimodal large language models (MLLMs) have made significant advancements in event-based vision. However, the comprehensive evaluation of these models' capabilities within a unified benchmark remains a crucial yet largely unexplored aspect. In this paper, we introduce **EventBench**, which presents an evaluation benchmark covering 8 diverse task metrics and a large-scale event stream dataset. Our work distinguishes existing event-based benchmarks in four key aspects: **i) Openness in accessibility**, releasing all raw event streams and task instructions across 8 evaluation metrics; **ii) Diversity in task coverage**, covering diverse understanding, recognition, and spatial reasoning tasks, enabling comprehensive evaluation of model capability; **iii) Integration in spatial dimensions**, pioneering the design of 3D spatial reasoning task for event-based MLLMs; **iv) Scale in data volume**, with an accompanying training set of over one million event–text pairs supporting large-scale training and evaluation. With EventBench, we evaluate state-of-the-art closed-source models such as GPT-5 and Gemini-2.5 Pro, leading open-source models including Qwen2.5-VL and InternVL3, as well as event-based MLLMs like EventGPT that directly process raw event inputs. Extensive evaluation reveals that current event-based MLLMs perform well in event stream understanding, yet they still struggle with fine-grained recognition and spatial reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.