What to Attend, Where to Act: Referent–Affordance Co-Evolution for Embodied Spatial Grounding
Abstract
Embodied spatial grounding refers to interpreting a language instruction and identifying the visual region for intended interaction. Existing methods broadly follow two paradigms: coordinate-generation methods produces discrete coordinate tokens as answers, whereas spatial-decoding methods predicts directly from visual representations to preserve spatial structure. However, both typically ground semantics into spatial predictions in a one-way manner, limiting reciprocal interaction between referent semantics and visual affordance evidence. To this end, we propose RACE, a Referent–Affordance Co-Evolution framework that couples referent semantics specifying what to attend to with an affordance state indicating where to act, jointly refining referent understanding and spatial localization. Specifically, referent semantics guide updates to the affordance state, which in turn feeds back visual evidence to refine referent understanding. To supervise guide this co-evolution, we unify heterogeneous grounding annotations as target distributions and combine distribution matching with a geometry-aware constraint penalizing spatially inconsistent predictions. Our model achieves 54.50% on RefSpatial-Bench and 74.00% on Where2Place while maintaining competitive general spatial-reasoning performance. These results suggest that referent–-affordance co-evolution provides an effective interface between task-relevant semantics and fine-grained spatial grounding, offering a general pathway from task-level reasoning to actionable spatial representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.