TraceEarth: From Semantic Matching to Spatial Tracing for Training-Free UHR Remote Sensing VQA
Abstract
Ultra-high-resolution (UHR) remote sensing visual question answering (VQA) requires capturing global context and local details within a limited visual token budget, making efficient acquisition of question-relevant local evidence critical.Existing methods improve UHR perception through visual token compression or selection, question-guided region retrieval, or iterative zooming and search. However, these methods mainly determine what visual content to retain or inspect based on semantic relevance or visual importance, while explicit spatial cues in the question are rarely converted into direct constraints on the search space. This can lead to redundant visual processing, localization errors, and missed critical evidence. We argue that remote sensing questions should not only identify ”which regions are relevant,” but directly indicate ”where the target is.” Converting spatial cues into executable visual search priors shifts broad semantic search toward efficient, precise target localization. We therefore propose TraceEarth, a training-free framework for spatially guided visual evidence acquisition in UHR remote sensing VQA. TraceEarth decouples spatial cues from full-question semantics, shifting visual evidence acquisition from semantic matching to spatial tracing. Building on this framework, we further propose a spatial cue guided evidence tracing method that combines image-referenced localization, object-referenced tracing, relation-preserving evidence construction, and candidate verification using frozen models and existing visual tools, progressively acquiring question-relevant evidence from global scenes to local targets. TraceEarth achieves 60.03% and 61.10% accuracy on the UHR benchmarks MME-RealWorld-RS and XLRS-Bench Lite, respectively, and 72.61% on the remote sensing benchmark LHRS-Bench. Consistent gains across base vision-language models demonstrate its effectiveness and generalizability without additional training. Our code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.