Thermo-VL: Extending Vision–Language Models to Thermal Infrared Perception
Abstract
Vision-language models (VLMs) often fail under low illumination because their visual grounding is learned predominantly from RGB imagery, whereas thermal infrared preserves complementary scene structure when visible cues degrade. We present **Thermo-VL**, a language-conditioned RGB-thermal VLM that augments a frozen Molmo-7B backbone with a trainable thermal encoder and a text-guided dual-attention fusion module. Given aligned RGB tokens, thermal tokens, and prompt embeddings, the fusion module conditions thermal features on both language and RGB context, then injects a gated residual into the frozen RGB stream so thermal evidence can be incorporated without disrupting Molmo's pretrained RGB-language interface. We train the model with the standard language-modeling objective together with auxiliary alignment and regularization losses that improve cross-modal grounding and reduce over-reliance on RGB. We also introduce a pixel-aligned RGB-thermal instruction-tuning dataset and Thermo-VL-Bench, a manually screened RGB-thermal VQA benchmark for low-light and cross-spectrum reasoning. Experiments show improvements in thermal-only reasoning and substantial gains in RGB+thermal reasoning, highlighting the value of prompt-conditioned RGB-thermal fusion. Data and code will be made publicly available after the review process.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.