Learning Task-Agnostic Representation for Multimodal Image Fusion
Abstract
Recent image fusion methods increasingly incorporate downstream tasks into joint optimisation, with the aim of injecting task-oriented semantic information into the fused image. However, compressing multimodal information into a single fused image inevitably imposes a fixed trade-off between the source modalities, making it difficult to accommodate the diverse information requirements of different scenes and downstream tasks. In this work, we explore an alternative paradigm. Rather than compressing complementary multimodal cues into a single fused representation, we aggregate source information under complementary fusion objectives to learn Task-Agnostic Representations (TAR). By allocating heterogeneous fusion cues that would otherwise compete within a shared output space, TAR preserves richer complementary information and provides a flexible intermediate representation for diverse downstream applications, while maintaining the same information budget as conventional image-fusion outputs. Importantly, TAR is learned without relying on task-specific annotations or downstream-task supervision, allowing the same representation to be directly exploited by different perception tasks. For conventional image fusion, a lightweight fusion head can further project TAR into a standard fused image, providing a flexible means of recovering visually faithful fusion results. Extensive experiments across multiple fusion and perception benchmarks demonstrate that TAR consistently improves downstream object detection and semantic segmentation performance while retaining competitive image-level fusion quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.