Grounding Beyond Tokens: Continuous Geometry from Visual Evidence
Abstract
Vision-language models increasingly formulate visual grounding as autoregressive generation of discrete coordinate tokens, unifying language understanding and spatial prediction within a single framework. However, standard token-level supervision treats coordinates as categorical labels without explicitly modeling their spatial ordering or geometric dependencies. Language-modeling objectives also supervise object-relevant visual evidence indirectly, leaving the correspondence among referring expressions, visual referents, and spatial predictions insufficiently constrained. In this work, we propose Locus, a generative framework for object detection and visual grounding that combines explicit visual evidence supervision with continuous coordinate prediction. Locus supervises the spatial distribution of phrase-conditioned visual attention, encouraging the alignment between linguistic representations and corresponding visual referents. It also interprets coordinate-token probabilities as distributions over an ordered spatial domain, enabling continuous coordinate estimation, distance-aware supervision, and box-level geometric optimization. Notably, Locus preserves the autoregressive generation paradigm and requires no additional detection modules or architectural changes. Experiments across detection and referring expression benchmarks demonstrate consistent gains in localization accuracy and state-of-the-art performance, supporting the effectiveness of combining visual evidence supervision with continuous geometric modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.