TFUSE: TEMPORAL FEEDBACK GUIDES CROSS- MODAL SELECTION FOR INFRARED–VISIBLE VIDEO FUSION
Abstract
Infrared–visible video fusion requires not only high-quality fusion at the individual-frame level, but also temporal consistency across consecutive frames. However, these two objectives may conflict with each other. To address this issue, we propose TFuse. At the objective level, we introduce a learnable loss generator. The generator performs undecimated Haar wavelet decomposition on infrared and visible clips, and predicts frame-adaptive kernels for the low- and high-frequency subbands to construct dynamic intensity losses. These losses are learned through online bilevel optimization inspired by the look-ahead principle of Nesterov acceleration. A support clip performs a differentiable virtual update; a query clip evaluates the updated model using a temporal-consistency objective, and the resulting temporal feedback optimizes the loss generator. At the representation level, we introduce Temporal Enhancement with Multi-frequency Propagation Operator (TEMPO). For low-frequency structural information, it jointly models temporal dependencies using clip-level global memory and bidirectional local memory. For high-frequency details, it employs confidence-gated temporal refinement to suppress unreliable temporal propagation. During cross-modal fusion, we introduce Modality-aware Information Routing and Aggregation (MIRA), which combines the basic fused features with the weighted sum of residuals from each modality to enable adaptive information selection. Through these designs, TFuse coordinates the competing objectives, improving temporal coherence while preserving high-quality fused imagery. Code is available at the supplementary materials.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.