When to Ground, Clarify, or Abstain: Evidence-Aware Selection and Grounding in Remote Sensing
Abstract
Multimodal large language models (MLLMs) have recently advanced remote sensing visual grounding (RSVG). However, existing RSVG methods typically assume that every query refers to one or more identifiable targets present in the image. This assumption forces models to produce a localization even when the intended target lacks sufficient visual support or the query does not contain enough information to distinguish it from similar objects. Existing work either handles these cases separately or maps both to abstention, rather than selecting between clarification and abstention according to the available evidence. To address this limitation, we formulate RSVG as a three-way selective grounding problem and propose EASG, an evidence-aware selection and grounding framework that requires the model to select Ground when one or more intended targets are sufficiently supported, Clarify when multiple candidates remain plausible, or Abstain when visual evidence is insufficient. Specifically, EASG uses a detector to extract coarse-grained candidate evidence and presents it to the MLLM as non-binding guidance, enabling the model to complete selection and grounding according to the available evidence. To support this task, we construct RS-EviBench, a benchmark comprising 68,328 samples that cover reliable grounding, referential ambiguity, and insufficient visual evidence. On the RS-EviBench test set, EASG achieves 96.89% decision accuracy, 58.83% mIoU, and 55.04% [email protected], improving mIoU and [email protected] by 8.01 and 13.36 percentage points, respectively, over the baseline, while experiments on VRSBench-VG and DIOR-RSVG further demonstrate its generalization to grounding-only RSVG benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.