acceptodds
Under review as a conference paper at ICLR 2027

Does RL Post-Training Ground Vision-Language Models? A Causal Measurement

Abstract

Reinforcement learning (RL) post-training raises the benchmark accuracy of vision-language models (VLMs), but it is unclear whether the gains come from better use of visual evidence or from stronger language priors. Accuracy cannot settle this, since many benchmark items are answerable without the image, and neither can attention, which need not reflect how the image is used. We introduce CARVE, a programmatic benchmark of solver-certified counterfactual edits: each edit changes one piece of visual evidence, certifies the new answer, and records the image region it occupies. We test whether answers follow the edit, whether they follow it when a printed value stays the same but moves to a different element, and whether the edited region's hidden states alone redirect the answer. Across four Qwen-VL backbones trained with GRPO, VPPO, and PRPO, post-training improves edit following modestly on three backbones, yet models that follow a numeral edit keep their original answer after 10.9–39.5% of edits that only move a label, against at most 1.7% of numeral edits, in all 16 configurations; post-training does not remove this failure. The gain is not reliably carried by the evidence: the edited region reproduces 43% of the behavioral gain in our 3B models (95% CI [18, 67]) but 96% in our 7B models ([66, 139]), and the 3B pattern persists with a frozen vision tower, locating the change in the decoder. Answer-changing evidence stays at visual positions rather than moving into the question, and attention to the image rises after training even where the causal role of visual states falls. Grounding claims for VLMs should therefore be tested against evidence whose effect on the answer is known, not inferred from accuracy or attention.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.