acceptodds
Under review as a conference paper at ICLR 2027

MF-RSVLM: Multi-Feature Fusion Vision-Language Model for Remote Sensing

Abstract

Large Vision-Language Models (VLMs) have demonstrated remarkable capabilities in the natural image domain, yet their adaptation to remote sensing (RS) imagery remains challenging. However, high-resolution and multi-scale inputs alone do not ensure effective integration of complementary visual details or sustained visual guidance during language-model processing. To address these challenges, we introduce MF-RSVLM, a multi-feature fusion vision–language model for RS understanding. Specifically, we extract complementary local features across spatial scales and encoder depths, organizing them into aligned detail representations alongside global context. To selectively integrate these details and sustain visual guidance, we use current visual states to guide adaptive fusion and regulate repeated injection at selected LLM layers, helping mitigate visual forgetting. To support unified multi-task learning, we curate a 293K-sample RS instruction corpus spanning six tasks. We evaluate MF-RSVLM across three benchmark families—image captioning, scene classification, and visual question answering. MF-RSVLM achieves state-of-the-art performance in the primary scene-classification and VQA comparisons, along with competitive image-captioning results. Complementary ablation and efficiency analyses examine the roles of detail sources, adaptive fusion, and repeated gated updates while quantifying their computational overhead. These analyses provide empirical guidance for multimodal model design through selective fusion and reuse of fine-grained visual information.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.