SatLAERS: A Resource-Efficient Remote Sensing VLM for Onboard Disaster Understanding
Abstract
Efficient remote sensing vision-language models (RS-VLMs) can enable onboard first-look disaster understanding from the latest available high-resolution optical observation, providing multi-granularity information for rapid response. Recent RS-VLMs have made substantial progress in unified multi-task understanding and model efficiency, yet disaster-oriented models remain comparatively limited and are often specialized to individual tasks. We introduce SatLAERS, a granularity-aware adaptive RS-VLM for resource-efficient adaptation and onboard inference. Built on a frozen vision encoder and compact language backbone, SatLAERS allocates multimodal computation according to task granularity and sample complexity, combining lightweight multi-scale visual adaptation and query-adaptive visual routing with sparse low-rank experts. A progressive three-stage pipeline proceeds from efficient vision-language alignment and multi-granularity instruction tuning to budget-aware reinforcement learning, jointly balancing task performance and computational cost. We further introduce SatDISA, a single-observation disaster benchmark with 22,862 high-resolution images and 316,698 vision-language instructions across six disaster types and four semantic granularities, spanning image-level understanding, region-level interpretation, pixel-level grounding, and disaster-oriented situational understanding. SatLAERS completes domain adaptation on 8 NVIDIA RTX 3090 GPUs in 49.3 GPU-hours, outperforms the strongest comparable-scale adapted baseline by 1.55 points, and achieves 16.99 s end-to-end latency with 4.461 GB peak RAM on Jetson Orin NX Super, supporting resource-constrained onboard deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.