acceptodds
Under review as a conference paper at ICLR 2027

What to Attend, Where to Act: Referent–Affordance Co-Evolution for Embodied Spatial Grounding

Abstract

Embodied spatial grounding refers to interpreting a language instruction and identifying the visual region for intended interaction. Existing methods broadly follow two paradigms: coordinate-generation methods produces discrete coordinate tokens as answers, whereas spatial-decoding methods predicts directly from visual representations to preserve spatial structure. However, both typically ground semantics into spatial predictions in a one-way manner, limiting reciprocal interaction between referent semantics and visual affordance evidence. To this end, we propose RACE, a Referent–Affordance Co-Evolution framework that couples referent semantics specifying what to attend to with an affordance state indicating where to act, jointly refining referent understanding and spatial localization. Specifically, referent semantics guide updates to the affordance state, which in turn feeds back visual evidence to refine referent understanding. To supervise guide this co-evolution, we unify heterogeneous grounding annotations as target distributions and combine distribution matching with a geometry-aware constraint penalizing spatially inconsistent predictions. Our model achieves 54.50% on RefSpatial-Bench and 74.00% on Where2Place while maintaining competitive general spatial-reasoning performance. These results suggest that referent–-affordance co-evolution provides an effective interface between task-relevant semantics and fine-grained spatial grounding, offering a general pathway from task-level reasoning to actionable spatial representations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.