acceptodds
Under review as a conference paper at ICLR 2027

RS-Ref: Benchmarking Reasoning in Remote Sensing Referring Expression Tasks

Abstract

Referring Expression Comprehension (REC) bridges natural language descriptions with specific image regions, serving as a fundamental capability for Vision-Language Models (VLMs). However, existing REC benchmarks in Remote Sensing (RS) frequently rely on simplistic expressions and lack sufficient distractors. This allows models to exploit single-attribute shortcuts rather than performing visual-language reasoning. In this paper, we investigate reasoning-intensive visual grounding in RS imagery, focusing on scenarios that demand multi-constraint logical reasoning amidst abundant same-category distractors. To this end, we introduce RS-Ref, a new benchmark specifically designed for multi-step remote sensing referring expression grounding. All images and expressions are strictly curated to ensure the presence of multiple same-category distractors and require compositional descriptions involving multiple constraints. Furthermore, we propose GeoREC, a two-stage visual chain-of-thought framework for referring expression. GeoREC introduces a grounded chain-of-thought mechanism to enable multi-step spatial reasoning, and subsequently employs GRPO-based reinforcement learning to optimize these reasoning trajectories. Extensive evaluations of 20 representative VLMs on RS-Ref, alongside validation on seven existing RS visual grounding benchmarks, demonstrate both the rigorous difficulty of RS-Ref and the state-of-the-art performance of our proposed framework.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.