acceptodds
Under review as a conference paper at ICLR 2027

SPDVF: Unified Infrared and Visible Video Fusion on Log-Euclidean Symmetric Positive Definite Manifold

Abstract

Existing infrared and visible video fusion methods treat spatial fusion and temporal modeling as separate stages in independent feature spaces. Temporal evolution is thus decoupled from fusion decisions, leaving flickering artifacts in the fused sequence. We observe that the second-order statistics of local features encode cross-modal complementarity and temporal dynamics at once, so a single descriptor can serve both roles. Such statistics take the form of symmetric positive definite (SPD) matrices, which reside on a curved manifold where linear operations are no longer valid. We propose SPDVF, a video fusion framework that describes local regions by SPD covariance descriptors and, under the Log-Euclidean metric that flattens this manifold, performs cross-modal alignment and temporal propagation in one shared tangent space without explicit motion estimation. Three components realize this design. Log-Euclidean Cross-Modal Attention aligns dual-modal second-order statistics through shifted-window cross-attention. The Basis-Kernel Bridge then converts the fused tokens into kernel mixing weights that refine Euclidean detail features. Log-Euclidean Mamba finally propagates temporal state over the same fused SPD tokens at a cost linear in clip length. Experiments on M3SVD, HDO, and VTMOT show that SPDVF improves spatial fidelity while maintaining temporal consistency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.