acceptodds
Under review as a conference paper at ICLR 2027

Parallel Instance Decoding for Language-Guided Localization in Aerial Images

Abstract

Earth observation (EO) localization is challenging due to large scene sizes, complex backgrounds, large variations in object scale and varying numbers of target objects. Existing remote-sensing vision-language models (RSVLMs) typically perform localization tasks, such as grounding, detection, and pointing by autoregressively generating bounding-box coordinates. As the object count increases, this step-by-step decoding becomes slow because boxes are produced one at a time and each prediction relies on earlier ones. To address this limitation, we introduce Parallel Instance Decoding (PID), a unified generative localization framework that separates instance identification from spatial localization. Given a textual reference (open vocabulary query), PID first generates a variable-length set of point addresses, with one address representing each matching object instance. It then predicts the bounding boxes for all identified instances in a single parallel decoding pass. To enable independent box prediction, we use an instance-isolated attention mask that allows each box to attend to the full image and its corresponding address while preventing attention to other boxes. Under the same training data and experimental setup, PID improves [email protected] over the direct-box baseline by 8.3 points on DIOR and 7.3 points on the DOTAv2 dataset. It also reduces complete-response latency by 22.4% and 43.6%, respectively. Moreover, PID performs favorably well against general-purpose VLMs and specialized RSVLMs on grounding and pointing tasks. We will publicly release our code and models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.