Same Reward, Different Ability: A Single-Variable Intervention Shows Task Formulation Governs What Spatial RLVR Learns
Abstract
Reinforcement learning with verifiable rewards (RLVR) is the standard recipe for improving spatial reasoning in vision–language models, and a growing diagnostic literature shows that such models often answer without looking. We ask the next question: what happens if the shortcut is removed before training? Nearly all point-level spatial benchmarks name the queried location as text coordinates, which the ground-plane relation makes informative about depth — a ridge regression that never sees the image beats a 7B vision–language model (VLM) on absolute depth. Holding the model, algorithm, reward, data volume, step count and the geometric distribution of training samples fixed, we vary only how the question refers to a pixel: drawn markers instead of coordinates. Across three paired seeds per arm, the marker formulation more than doubles metric accuracy and lifts every depth bucket off zero, while on these backbones the coordinate formulation leaves the hardest bucket at exactly zero under both RLVR and supervised fine-tuning. A decomposition arm shows the effect is the sum of two sub-interventions rather than one. The learned ability survives an adversarial set of off-ground points where the shortcut is void by construction, but does not transfer indoors, and ordinal accuracy stays below chance on a matched adversarial set — training removes the metric shortcut, not the ordinal position prior. Repeating the audit on two further backbones, including one from a different vendor, reproduces the diagnosis. Repeating the intervention shows that backbone capacity decides whether the formulation governs existence or merely degree — and that on the second vendor's model the effect does not appear at all, because that backbone learns no measurable metric ability here to begin with.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.