acceptodds
Under review as a conference paper at ICLR 2027

RETDet: Degradation-Aware Text as a Detachable Prior for RGB-Event Object Detection

Abstract

By fusing RGB images and event data, multimodal object detection improves accuracy and robustness in complex scenes, yet attention-dominated fusion modules incur high computational cost, and under degraded conditions, the representation capability of both modalities declines. High-level semantic priors describing the degradation itself could compensate for this loss, but querying a multimodal large language model (MLLM) at every inference step is prohibitively expensive. We propose RETDet, an RGB-Event-Text detection framework in which degradation-aware text acts as a detachable prior. During training, MLLM-generated degradation descriptions are mapped into layer-wise modulation parameters and randomly replaced by a learned null embedding, so that the detector internalizes degradation priors into its visual pathway while remaining able to consume explicit descriptions. The trained model supports two inference modes: a prior mode taking only RGB and event inputs, invoking neither an MLLM nor a text encoder at deployment, and a conditioned mode accepting scene descriptions for higher accuracy. Without any textual input at inference, prior mode already outperforms SOTA RGB-Event methods on DSEC, PKU-DDD17-Car, TUMTraf, and EventKITTI, and conditioned mode yields further gains concentrated on the most degraded subsets, decoupling the semantic benefit of MLLM supervision from its inference cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.