FREST: Frequency Representation Enhancement for Spatiotemporal Robustness in Infrared and Visible Video Fusion
Abstract
Compared with single-frame image fusion, infrared and visible video fusion better reflects real-world sensing scenarios and supports continuous scene perception. However, existing spatial-domain methods struggle to separate degradation interference from informative structures and distinguish abnormal temporal fluctuations from valid motion, resulting in limited robustness to complex degradations and compromised temporal coherence. To address these issues, we propose FREST, which leverages frequency representations to suppress interference from degradation and abnormal temporal fluctuations while preserving informative content and valid motion. Specifically, it comprises two components: degradation-robust spatial-frequency modeling and frequency-aware spatiotemporal modeling. On the one hand, to suppress degradation-sensitive responses while preserving scene structures, we develop a degradation-robust spatial-frequency modeling framework to exploit the distinct spectral characteristics of degradation interference and scene structures. Specifically, we develop an adaptive frequency enhancement module and a spatial-frequency collaborative modeling strategy for dynamic modulation and effective cross-modal interaction. On the other hand, given that abnormal flickering and valid motion exhibit distinct response patterns in the joint spatiotemporal frequency domain, we construct a frequency domain temporal consistency module to model spectral changes across consecutive frames, thereby reducing unstable temporal fluctuations while preserving meaningful motion. Extensive experiments demonstrate that FREST produces more temporally coherent fused videos than state-of-the-art methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.