Making Localization Thinkable: Grounding Actions as Thought Primitives
Abstract
Reasoning has become fundamental to modern Vision-Language Models (VLMs), extending visual understanding beyond direct prediction. In visual grounding, however, reasoning remains largely semantic, leaving precise localization outside the reasoning process. We term this gap Reasoning-to-Coordinate Collapse: VLMs can identify the correct target yet still predict inaccurate coordinates because localization is not explicitly modeled during reasoning. We present ThinkAxis, a paradigm that brings localization directly into the reasoning process. ThinkAxis represents localization updates as executable grounding primitives and carries the evolving localization throughout the reasoning trajectory, enabling VLMs to reason over and act on coordinates alongside semantic interpretation. To optimize this localization reasoning process, we further investigate the visual evidence behind localization decisions and find that accurate coordinate evolution increasingly relies on boundary-relevant evidence. Building on this observation, we formulate a trajectory optimization algorithm, ERPO, which couples geometric progress with counterfactual evidence attribution. Extensive experiments across visual grounding tasks show the effectiveness of ThinkAxis, with particularly notable improvements in precise localization under strict localization criteria. These results highlight the potential of localization-space reasoning for broader multimodal systems where precise spatial grounding is critical.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.