acceptodds
Under review as a conference paper at ICLR 2027

Visual Reasoning with Chain-of-Coordinates

Abstract

We introduce Chain-of-Coordinates (CoC), a representation for visual reasoning in which vision-language models (VLMs) autoregressively generate sequences of image coordinates. We provide a theoretical analysis of CoC, covering its representability, learnability, geometric approximation error, and computational advantage over direct answering on a family of connectivity problems. We introduce a two-stage reinforcement learning (RL) framework with a Fréchet-based geometric reward. Across six tasks, CoC improves curve tracing and consistently outperforms direct-answer training on reasoning tasks, with further gains from RL. Additional analyses show that CoC is most beneficial when connectivity cannot be inferred from direct visual cues, while the effectiveness of CoC prompting depends on a model's underlying tracing ability. Together, these results establish CoC as a paradigm for connecting visual perception and reasoning through explicit coordinate sequences.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.