acceptodds
Under review as a conference paper at ICLR 2027

Region-Level Credit Assignment from On-Policy Rollouts in Multimodal Reinforcement Post-Training

Abstract

Multimodal reinforcement post-training can improve a response's reward without resolving the visual mistakes within it. A shared outcome advantage can reinforce incorrect regional selections alongside correct ones, while omitted objects receive no explicit target in the sampled response. This leaves a gap between optimizing response quality and learning accurate, complete visual grounding. We preent MOVE (Model’s own Visual Errors), a framework that turns the policy's own regional errors into corrective supervision. MoVE contrasts corrected and error-preserving regional conditions to refine credit within a sampled response, and distills corrected targets to supervise objects absent from it. These complementary signals address both incorrect selections and omissions, tying training feedback to specific, verifiable regional mistakes. MoVE provides a unified framework for region-aware multimodal post-training, extending GRPO, DAPO, and OPD with targeted visual correction. Inference still uses only the original image and query. Across five grounding benchmarks, MoVE improves average Micro-F1 in matched GRPO and DAPO comparisons at both 2B and 4B, including an increase from to over DAPO with Qwen3-VL-2B. Corrected-region training benefits three model families, while the extension to visual question answering improves mean accuracy by relative to the base models across four fine-grained benchmarks and two model scales.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.