Sparse Frames, Causal Graphs: Long Event-Stream Anomaly Reasoning with Vision-Language Models
Abstract
RGB sensing can lose critical anomaly evidence under fast motion or adverse illumination, while event cameras preserve high-dynamic-range, microsecond-scale changes. Most event-language methods nevertheless aggregate events into frame-like inputs, obscuring the temporal detail that makes event sensing valuable. We propose EFG-VLM, an offline hybrid event-frame/event-graph framework for long-stream open-world anomaly understanding. A category-agnostic locator establishes shared anchors where sparse event frames provide scene semantics and bounded raw-event graphs preserve complementary spatial structure and causal motion. Hierarchical fusion combines landmark, proposal, and video context and exposes structured evidence to a pretrained multimodal language model for anomaly judgement, proposal-conditioned temporal localization, and description. Because localization is category-agnostic and final interpretation is language-generated rather than closed-set classification, the same interface can describe interactions outside benchmark categories. On large-scale v2e-simulated event streams derived from HIVAU-70K, EFG-VLM substantially improves event-source recognition and language description while remaining robust across clip durations. Native RGB–event cases further illustrate transfer to high-dynamic-range interactions beyond the benchmark categories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.