CEDVS-Bench: Speak or Stay Silent? A Benchmark for Continuous Event Detection in Video Streams
Abstract
Video assistants must recognize events while they unfold, decide when evidence justifies a report, and remain silent when a requested event never occurs. Retrospective localization and answers at selected query times only partially capture this requirement. We introduce CEDVS-Bench, a unified benchmark for continuous detection of language-defined events in video streams. Candidate descriptions are specified before observation; models make causal decisions every second, without access to future frames or event boundaries. The benchmark contains 2,840 videos, 15,983 positive event instances, and 66,158 absent candidates across short, medium, and long video settings. Our evaluation jointly measures onset localization, decision timeliness, processing-aware response availability, and rejection of absent events. Experiments on both window-based VLMs and native streaming models show that accurate localization does not necessarily lead to timely reporting, while semantically similar absent events remain difficult to reject, especially over longer observation horizons. CEDVS-Bench provides a unified testbed for accurate, timely, and reliable continuous event monitoring in video streams.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.