acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Infrared-Visible Image Fusion via Semantic Diagnosis and Reward-Driven Optimization

Abstract

Infrared and visible (IR-VIS) image fusion aims to integrate thermal target saliency from IR images and textural details from VIS images into a single fused image. Existing methods mainly rely on low-level visual fidelity objectives, emphasizing image sharpness and structural consistency while overlooking semantic preservation and task-relevant cues. Moreover, the lack of ground-truth fusion labels and scene-dependent trade-offs among thermal saliency, texture details, structural consistency, and artifact suppression make fixed loss functions inadequate for adaptive optimization. To address these limitations, we propose a semantic-diagnosis- and reward-driven framework that combines high-level semantic guidance with candidate-based reward optimization. Specifically, we introduce a frozen DINOv3-ConvNeXt to provide hierarchical visual priors and global visual representations, together with a lightweight adaptation branch for local texture and detail enhancement. Building on these representations, we leverage a frozen Qwen3-VL-Embedding to perform coarse-to-fine semantic diagnosis on intermediate fusion states and progressively guide their refinement. To further enable adaptive optimization without fusion labels, we generate multiple fusion candidates through perturbation sampling and evaluate their relative quality using multidimensional rewards. We then employ an advantage-weighted objective to optimize the candidate generation policy, thereby learning fusion policies through candidate comparison and reward optimization. Experimental results demonstrate that our method consistently outperforms existing approaches in both quantitative metrics and visual quality across multiple public benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.