Region-Level Credit Assignment from On-Policy Rollouts in Multimodal Reinforcement Post-Training
Abstract
Multimodal reinforcement post-training can improve a response's reward without resolving the visual mistakes within it. A shared outcome advantage can reinforce incorrect regional selections alongside correct ones, while omitted objects receive no explicit target in the sampled response. This leaves a gap between optimizing response quality and learning accurate, complete visual grounding. We preent MOVE (Model’s own Visual Errors), a framework that turns the policy's own regional errors into corrective supervision. MoVE contrasts corrected and error-preserving regional conditions to refine credit within a sampled response, and distills corrected targets to supervise objects absent from it. These complementary signals address both incorrect selections and omissions, tying training feedback to specific, verifiable regional mistakes. MoVE provides a unified framework for region-aware multimodal post-training, extending GRPO, DAPO, and OPD with targeted visual correction. Inference still uses only the original image and query. Across five grounding benchmarks, MoVE improves average Micro-F1 in matched GRPO and DAPO comparisons at both 2B and 4B, including an increase from to over DAPO with Qwen3-VL-2B. Corrected-region training benefits three model families, while the extension to visual question answering improves mean accuracy by relative to the base models across four fine-grained benchmarks and two model scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.