TriReward-Fusion: Unified Vision-Language Reward Learning for Task-Oriented Infrared–Visible Image Fusion
Abstract
Infrared and visible image fusion aims to combine the thermal saliency of infrared images with the structural details of visible images for robust high-level visual perception. However, most existing fusion methods are optimized using handcrafted no-reference pixel-level objectives or supervised by a single downstream task, making it difficult to jointly preserve visual fidelity, semantic integrity, and object-level reliability. Moreover, image fusion, semantic segmentation, and object detection pursue distinct and sometimes conflicting objectives, and a simple weighted combination of their losses often fails to provide unified and semantically meaningful feedback. To address this problem, we propose TriReward-Fusion, a BLIP-guided vision-language reward learning framework for task-oriented infrared-visible image fusion. The proposed framework is trained on MSRS by jointly using image-pair no-reference fusion objectives, semantic segmentation annotations, and our additionally annotated detection labels. It projects fusion, segmentation, and detection outputs into a shared vision-language semantic state space, where perceptual information preservation, semantic region completeness, and instance-level detection reliability are transformed into learnable multi-task reward signals. These rewards provide unified semantic feedback to guide the fusion process and reconcile pixel-level reconstruction, region-level parsing, and object-level localization objectives. Experiments evaluate fusion quality on MSRS, M3FD, and LLVIP, semantic segmentation on MSRS, and object detection on M3FD. The results show that TriReward-Fusion improves fused image quality and enhances downstream segmentation and detection performance, demonstrating the effectiveness of vision-language reward learning for unified multi-task infrared-visible image fusion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.