acceptodds
Under review as a conference paper at ICLR 2027

GEOEPS: TRAINING-FREE EVIDENCE PRESENTATION AND SELECTION FOR ULTRA-HIGH-RESOLUTION REMOTE SENSING UNDERSTANDING

Abstract

Ultra-high-resolution (UHR) remote-sensing (RS) imagery poses a fundamental challenge for vision-language models (VLMs): task-relevant evidence may occupy only a tiny fraction of a large scene, yet making that evidence more accessible does not ensure its effective use in the final answer. We identify this discrepancy as an evidence-to-decision gap, which is often observed in visual question answering (VQA) tasks for UHR-RS imagery. Through controlled comparisons, we find that how a referenced region is presented matters beyond simply locating it. To address this gap, we introduce GeoEPS, a training-free framework that coordinates evidence presentation and selection with frozen VLMs. When the input provides an explicit spatial cue, Evidence-aware Presentation (EAP) uses it as an evidence prior and combines local context, textual alignment, unfilled marking, and query-conditioned presentation. For inputs without an explicit spatial cue, Competitive Evidence Selection (CES) searches local evidence candidates and compares them under competing answer hypotheses. On XLRS-Bench and RSHR-Bench, our GeoEPS substantially improves Qwen3-VL-8B's average task accuracy from 48.8% to 61.0% and from 41.3% to 53.7%, respectively. Extensive experiments further demonstrate that these gains can extend to other frozen general-purpose and RS-specific VLMs without parameter updates or UHR-specific retraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.