H-GRPO: PERMUTATION-INVARIANT REINFORCE- MENT LEARNING FOR IMPROVING SPATIAL GROUNDED REASONING
Abstract
Vision-language models (VLMs) have achieved strong performance on visual reasoning tasks, yet a correct final answer does not necessarily imply that the model followed a correct or visually grounded reasoning process. We study how VLMs can be trained to produce explicit intermediate reasoning steps that are both semantically meaningful and grounded in visual evidence. We represent each reasoning step as a structured triplet consisting of a sub-question, sub-answer, and supporting bounding box, enabling the reasoning process to be evaluated along both semantic and spatial dimensions. To optimize this representation, we introduce Hungarian-GRPO (H-GRPO), a reinforcement-learning framework that aligns predicted and reference reasoning steps using permutation-invariant Hungarian matching, avoiding the assumption that valid reasoning steps must follow a fixed ordering. Experiments show that optimizing grounded reasoning improves both intermediate reasoning quality and answer accuracy, with H-GRPO variant achieving the strongest overall grounding and decomposition performance. Our analyses show that grounding quality and answer correctness are related but distinct, while counterfactual evidence masking confirms that the localized regions contain information important to the model's predictions. These findings motivate evaluating VLMs beyond answer accuracy and demonstrate the value of explicit, spatially grounded reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.