acceptodds
Under review as a conference paper at ICLR 2027

EfficientViS: Adaptive Parallel Visual Search for High-Resolution Remote Sensing Visual Grounding

Abstract

High-resolution remote sensing visual grounding (RSVG) requires both global spatial context and fine-grained evidence for small targets. Full-scene inference preserves context but loses local details after resizing, whereas iterative visual search recovers such details through sequential crop-and-ground calls at substantial computational cost. We propose EfficientViS, an adaptive parallel framework for efficient high-resolution RSVG. EfficientViS consists of two components: 1) One-shot Complementary Region Planning (OCRP), which predicts diverse candidate regions from a single global image–text representation before local grounding; and 2) Adaptive Candidate Parallel Inference (ACPI), which selects a query-adaptive observation budget and packs selected crops into a single specialist forward pass for efficient grounding. We further introduce a two-stage training strategy: Stage I trains a crop-robust specialist grounder to improve local grounding, while Stage II applies supervised-to-PPO training to optimize the accuracy–cost trade-off. Extensive experiments on three benchmarks (RSVG-HR, AVVG, and XLRS-Bench) demonstrate that EfficientViS delivers 4.3–4.8× higher throughput than the corresponding full-scene specialist baselines and more than 14× higher throughput than sequential search.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.