Fine-Grained Geometric Visual Perception: Recovering Precise Positions, Boundaries, and Trajectories
Abstract
Vision–language models (VLMs) can recognize geometric objects and coarse relations, yet they remain unreliable when asked to recover precise positions, endpoints, boundaries, and continuous trajectories in a common canvas co- ordinate system. We formulate this capability as fine-grained geometric vi- sual perception (FGGVP) and introduce FGGVP-22K, a dataset containing approximately 20,000 synthetic samples and 1,500 real samples. The data cover explicit primitive localization, semantically specified components, and derived targets determined by geometric relations. We fine-tune Qwen3- VL-4B with native coordinate tokens and multitask supervision, combining autoregressive answer loss with ordinal-coordinate, expected pixel-distance, and coordinate-point CVaR losses. Experiments on synthetic, real, and ex- ternal geometric datasets show substantial improvements in fine-grained localization over the base model and stronger Image2Code reconstruction. The results indicate that targeted supervision of positions, boundaries, and trajectories improves geometric visual perception and transfers to down- stream reconstruction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.