acceptodds
Under review as a conference paper at ICLR 2027

GaIF: Content-Adaptive Patch–Channel Gating for Image Fusion

Abstract

Infrared–visible image fusion requires spatially adaptive modality integration, as various regions demand different balances between thermal saliency and visible texture. However, existing methods often operate at an overly coarse fusion granularity: CNN-based autoencoders merge modality streams before fine-grained routing, while scalar-gated transformers apply a single fusion coefficient uniformly across tokens. This causes a mismatch with the inherently nature of fusion, where modality contributions should vary across patches and feature channels. To address this limitation, we propose GaIF with Patch-Conditioned Modality Routing (PCMR). Bidirectional cross-attention first introduces cross-modal context into into each token, after which a lightweight MLP predicts a gate to enable patch–channel-specific routing at each layer. A confidence-weighted compositor further handles uncertain regions, while an intra-sequence depthwise mixer improves local spatial coherence. Trained only on LLVIP and synthetic multi-focus pairs, GaIF generalizes across IVIF, MEIF, remote sensing, and NIR–VIS fusion tasks, achieving better results than GIFNet except for SCD on Lytro, MFI-WHU, and Harvard datasets. Ablation studies validate each component independently. These results demonstrate the effectiveness and generalizability of fine-grained patch–channel modality routing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.