Relation Matters: Learning Local Relations for Binary Infrared-Visible Image Fusion
Abstract
Infrared–Visible Image Fusion (IVIF) integrates thermal targets with spatial details, but full-precision models require substantial computation and storage. Binarization reduces these costs, yet binary IVIF faces three connected challenges: binary representation, where binarization can obscure local structural differences; feature propagation, where binary processing can degrade the magnitude information needed for reconstruction; and capacity-compatible training, where the binary network must learn complementary fusion patterns within its limited representational capacity. To address these challenges, we propose SDRTNet, a binary IVIF framework that jointly addresses binary representation, feature propagation, and capacity-compatible training. Sliding Directional Relation Transform (SDRT) makes local directional differences explicit by decomposing overlapping patches into a base component and four directional relation components for separate binary processing. Local Statistical Alignment (LSA) then uses the local means and variances of raw SDRT features to calibrate binary convolution outputs before spatial reconstruction. During training, Binary Vector-Quantized Knowledge Distillation (BVQKD) organizes a full-precision teacher's fusion knowledge into a shared codebook and guides the binary student through codeword-assignment distributions, without adding inference cost. Experiments on four fusion benchmarks and two downstream tasks show competitive fusion performance with reduced computational and storage costs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.