TF-DETR: Task-Factor Representation for RGB-X Object Detection
Abstract
RGB-X object detection typically represents RGB and auxiliary modalities as parallel feature streams and focuses on how their representations should be fused. However, given an already strong RGB representation, we observe that the additional task utility provided by X varies substantially across sensing modalities and detection tasks.Motivated by this observation, we propose TF-DETR (Task-Factor DETR), which preserves RGB as the primary semantic stream while representing the auxiliary modality through Task Factors that combine explicit detection-oriented states with flexible sensor-dependent latent evidence. To exploit these factors effectively, TF-DETR establishes cross-modal correspondence through global-anchored multi-scale task alignment, and incorporates aligned auxiliary evidence into the RGB encoder through reliability-aware query guidance and bounded low-rank feature refinement. This design allows X to guide and refine RGB representations without maintaining a second full semantic stream.Experiments on RGB-Thermal and RGB-Event object detection benchmarks demonstrate state-of-the-art performance with a lightweight auxiliary pathway. Code will be released publicly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.