acceptodds
Under review as a conference paper at ICLR 2027

Reasoning in the Dark: Event-RGB Adaptation for Low-Light Video MLLMs

Abstract

Video multimodal large language models (video MLLMs) rely heavily on RGB appearance cues and degrade sharply in low light, where texture, color, and object boundaries are lost. Event cameras provide complementary motion and temporal information under poor illumination, but integrating event streams into language-based reasoning is bottlenecked by scarce paired event-language supervision, especially under low light. We propose LITE-Event, a sample-efficient framework for low-light video reasoning that fuses event and RGB streams. A lightweight visual-language interface combines modality-specific projectors with gated fusion. A support-conditioned hypernetwork then generates LoRA-style residual updates from a small non-target support set, adapting the interface without updating the frozen encoders or LLM. We also introduce three resources for training and evaluation: NExT-QA-Dark, a large-scale synthetic low-light RGB-event dataset for pretraining, EventLang-Dark, a curated event-language dataset for adaptation, and RealDark-Bench, a real-world low-light RGB-event reasoning benchmark. Trained predominantly on synthetic data, LITE-Event reaches 85.0% on RealDark-Bench, outperforming the strongest enhancement-then-reason baseline by 5.2 points and surpassing both RGB-only and event-only video MLLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.