AnyPoint: Unpacking Cross-View Understanding Granularity in Vision-Language Models
Abstract
Fine-grained cross-view understanding requires Vision-Language Models (VLMs) to localize arbitrary pixel locations across changes in viewpoint and appearance. This capability is fundamental to spatial reasoning and embodied AI. Existing benchmarks rely on discrete candidate selection or isolated pixel queries. However, discrete candidate selection limits localization precision, while isolated pixel queries yield query-dependent scores and fail to assess global understanding. To address these limitations, we introduce AnyPoint-Bench, the first benchmark to formulate fine-grained cross-view understanding as the cross-view localization of arbitrary pixels. Given an image pair, a VLM jointly localizes an arbitrary set of reference-view pixels by predicting their coordinates in the target view. This unified protocol evaluates both fine-grained localization and global understanding. AnyPoint-Bench covers diverse scenes, viewpoint and appearance changes, and query configurations. Our evaluation reveals a pronounced gap between proprietary and open-source VLMs. GPT-5.6-Sol achieves 36.07% accuracy, compared with 17.19% for Qwen3.8. To probe the potential of the VLM paradigm, we develop CorrVLM, a cross-view specialist trained with vocabulary repurposing and coarse-to-fine learning. Despite using substantially fewer image pairs than specialized visual-geometric models, CorrVLM matches their performance at moderate thresholds and surpasses GPT-5.6-Sol by 31.28 percentage points. These results demonstrate that strong fine-grained cross-view understanding is attainable within the VLM paradigm, yet remains largely unrealized in current general-purpose VLMs. AnyPoint-Bench offers a new lens for evaluating this capability and advancing fine-grained cross-view understanding in VLMs. Our benchmark, code, and models will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.