Seeing the Whole Scene: Scaling Vision-Language Models for Ultra-High-Resolution Earth Observation
Abstract
Ultra-high-resolution (UHR) Earth observation imagery enables fine-grained reasoning over large areas, but rapidly growing visual sequences make full-attention remote sensing vision-language models (RS-VLMs) difficult to scale. Existing UHR RS-VLMs typically control cost through multiscale tiling, adaptive region selection, or token pruning and merging. These strategies bound computation by reducing the visual sequence, requiring increasingly aggressive visual compression as image size grows and potentially discarding fine spatial detail or query-relevant information. We introduce SourceScale-VL, a post-training-converted RS-VLM that preserves the complete post-encoder visual sequence and scales it through ring-style sequence parallelism, compressing visual interactions rather than visual content. SourceScale-VL replaces quadratic visual propagation with linear attention and introduces a magnitude-aware Referential Spectrum Residual to selectively retain Exact text-to-visual interactions where linear attention lacks sufficient source-response capacity. To preserve source-specific behavior during conversion, we further introduce Source-Factored On-Policy Distillation, which aligns source-wise output contributions on student-generated states. On the held-out conversion-evaluation split, SourceScale-VL reaches 57.46% accuracy, outperforming the same-budget RSR-Hybrid by 0.43 percentage points while reducing SCE from 0.063 to 0.041. On XLRS-Bench-lite and LRS-VQA, it achieves 54.64% and 33.58% aggregate accuracy, slightly surpassing Teacher (Exact) by 0.14 and 0.18 percentage points, respectively. With four A800 GPUs, SourceScale-VL supports longer visual sequences than sequence-parallel Exact attention while reducing peak memory by 13.8% and TTFT by 15.5%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.