SPDVF: Unified Infrared and Visible Video Fusion on Log-Euclidean Symmetric Positive Definite Manifold
Abstract
Existing infrared and visible video fusion methods treat spatial fusion and temporal modeling as separate stages in independent feature spaces. Temporal evolution is thus decoupled from fusion decisions, leaving flickering artifacts in the fused sequence. We observe that the second-order statistics of local features encode cross-modal complementarity and temporal dynamics at once, so a single descriptor can serve both roles. Such statistics take the form of symmetric positive definite (SPD) matrices, which reside on a curved manifold where linear operations are no longer valid. We propose SPDVF, a video fusion framework that describes local regions by SPD covariance descriptors and, under the Log-Euclidean metric that flattens this manifold, performs cross-modal alignment and temporal propagation in one shared tangent space without explicit motion estimation. Three components realize this design. Log-Euclidean Cross-Modal Attention aligns dual-modal second-order statistics through shifted-window cross-attention. The Basis-Kernel Bridge then converts the fused tokens into kernel mixing weights that refine Euclidean detail features. Log-Euclidean Mamba finally propagates temporal state over the same fused SPD tokens at a cost linear in clip length. Experiments on M3SVD, HDO, and VTMOT show that SPDVF improves spatial fidelity while maintaining temporal consistency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.