Assessing Infrared-Visible Fusion Towards Machine Perception and Understanding Preferences: Benchmark, Dataset, and Baseline
Abstract
Despite substantial progress in infrared-visible image fusion (IVIF) over the past decades, fused-image evaluation still relies heavily on metrics adapted from generic image quality assessment, which primarily cater to human perceptual preference. Machine-centric evaluation, by contrast, remains underexplored due to the lack of dedicated benchmarks. To address this limitation, we formalize machine preference as the utility of fused images for representative downstream perception and understanding tasks. We then introduce FuM-Score, which jointly quantifies such preference by aggregating fusion gains relative to the infrared and visible sources across these tasks under standardized evaluation protocols. Based on these, we subsequently establish FuM-18K, to our knowledge, the first large-scale dataset for machine-centric assessment in IVIF, comprising over 18,000 images spanning diverse scenes and representative fusion models, along with fine-grained downstream task annotations. Finally, an effective FuM-Net is built to leverage fusion features together with auxiliary source references, achieving automatic FuM-Score prediction for quality assessment. Extensive experiments validate the efficacy of our benchmark in capturing machine preference and demonstrate the superiority of FuM-Net over recent state-of-the-art methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.