ART: From Common Semantics to Unique Details via Decoupled Feature Interaction for RGBT Detection
Abstract
RGB–thermal (RGBT) object detection is essential for around-the-clock perception. However, many existing detectors implicitly couple fusion with interaction—pushing aggressive cross-modal mixing early—which can prematurely collapse modality diversity and suppress complementary cues. We argue that effective interaction should be built on the right factors: shared semantics should be aligned, while modality-specific details should be preserved. Based on this insight, we present ART, a decoupled feature interaction framework that follows Decouple Align Interact. ART first disentangles features into cross-modal common semantics and modality-specific unique details using a modality-agnostic shared projection head and modality-dependent heads. It then stabilizes the shared semantic subspace with a correlation-based alignment objective that encourages an identity-like cross-modal channel correlation—strengthening matched-channel semantic consistency while reducing off-diagonal leakage and redundancy. Finally, ART performs cross-modal interaction separately in the common and unique spaces, enabling semantic agreement without erasing fine-grained complementary information. Experiments on public RGBT benchmarks demonstrate that ART achieves state-of-the-art performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.