acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Task-Oriented Image Fusion as Probabilistic Inference

Abstract

Although task-oriented image fusion has achieved promising downstream performance by incorporating task supervision, existing methods typically learn a direct mapping from multimodal inputs to a single fused image, with task objectives serving mainly as optimization signals. Such a formulation leaves the underlying conditional distribution of task-relevant fusion results unmodeled, making it difficult to capture the distributional structure and statistical regularities of multimodal integration for downstream tasks. To address this limitation, we rethink task-oriented image fusion as probabilistic inference, modeling the fused image as a conditional distribution jointly determined by source modalities and downstream task information. We introduce a latent fusion variable that captures both multimodal fusion cues and task-relevant semantics, which is sampled and decoded to generate fused images. During training, task annotations are used to construct a task-informed posterior, whose knowledge is transferred to a source-conditioned prior for label-free inference. We further align source modal and fused image task predictions to preserve task-relevant information. Experiments on infrared-visible fusion with semantic segmentation and object detection demonstrate consistent improvements in both fusion quality and downstream performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.