UAV-VIVID: A Large-Scale UAV Infrared–Visible Video Dataset and a Text-Guided Frame-Level Fusion Baseline
Abstract
Infrared–visible (IR–VI) fusion is commonly developed on independent, well-aligned image pairs, leaving its behavior on continuous UAV streams insufficiently characterized. We introduce UAV-VIVID, a UAV-acquired IR–VI video dataset containing 375 temporally continuous scene segments and 190,978 paired frames. Its 96,087 registration-processed lower-altitude pairs define the fusion track evaluated in this work, while 79,992 higher-altitude pairs and 14,899 unregistered lower-altitude pairs are released as auxiliary resources. The dataset preserves temporal order, adopts scene-disjoint splits, and provides structured textual descriptions. To establish a controlled frame-level reference for this video setting, we further present the Residual Text-Guided Fusion Network (RT-FuseNet), which processes each IR–VI pair independently without neighboring frames or temporal state. Adaptive Residual Balancing regulates cross-modal interactions, while Text-Affine and Gated Modulation injects language guidance through channel-wise affine parameters and spatial gates. Experiments across multiple datasets evaluate frame-wise fusion quality, sequence-level output behavior, text conditioning, downstream utility, and the clean-input regularization effect of training-time dual-modal PGD augmentation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.