acceptodds
Under review as a conference paper at ICLR 2027

Grounding Beyond Tokens: Continuous Geometry from Visual Evidence

Abstract

Vision-language models increasingly formulate visual grounding as autoregressive generation of discrete coordinate tokens, unifying language understanding and spatial prediction within a single framework. However, standard token-level supervision treats coordinates as categorical labels without explicitly modeling their spatial ordering or geometric dependencies. Language-modeling objectives also supervise object-relevant visual evidence indirectly, leaving the correspondence among referring expressions, visual referents, and spatial predictions insufficiently constrained. In this work, we propose Locus, a generative framework for object detection and visual grounding that combines explicit visual evidence supervision with continuous coordinate prediction. Locus supervises the spatial distribution of phrase-conditioned visual attention, encouraging the alignment between linguistic representations and corresponding visual referents. It also interprets coordinate-token probabilities as distributions over an ordered spatial domain, enabling continuous coordinate estimation, distance-aware supervision, and box-level geometric optimization. Notably, Locus preserves the autoregressive generation paradigm and requires no additional detection modules or architectural changes. Experiments across detection and referring expression benchmarks demonstrate consistent gains in localization accuracy and state-of-the-art performance, supporting the effectiveness of combining visual evidence supervision with continuous geometric modeling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.