Visual Reasoning with Chain-of-Coordinates
Abstract
We introduce Chain-of-Coordinates (CoC), a representation for visual reasoning in which vision-language models (VLMs) autoregressively generate sequences of image coordinates. We provide a theoretical analysis of CoC, covering its representability, learnability, geometric approximation error, and computational advantage over direct answering on a family of connectivity problems. We introduce a two-stage reinforcement learning (RL) framework with a Fréchet-based geometric reward. Across six tasks, CoC improves curve tracing and consistently outperforms direct-answer training on reasoning tasks, with further gains from RL. Additional analyses show that CoC is most beneficial when connectivity cannot be inferred from direct visual cues, while the effectiveness of CoC prompting depends on a model's underlying tracing ability. Together, these results establish CoC as a paradigm for connecting visual perception and reasoning through explicit coordinate sequences.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.