acceptodds
Under review as a conference paper at ICLR 2027

Pixel-OPSD: Beyond Scalar Rewards for Pixel-Level MLLMs via On-Policy Self-Distillation

Abstract

Pixel-level multimodal large language models must preserve a region's meaning in language and its extent in pixels. Scalar outcome rewards assess the final prediction but do not specify how to correct its semantic and spatial errors. We introduce Pixel-OPSD, which shifts pixel-level learning from scalar-reward reinforcement learning to on-policy self-distillation (OPSD). Our central idea is to make evidence-conditioned self-teaching responsive to both the semantic state of a generated trajectory and the spatial effects of its mask codes. Reliability routing adapts semantic correction to the sampled description through verified replacement, on-policy prefix refinement, or retention. Evidence-Guided Credit Assignment (EGCA) grounds mask-code credit in changes between decoded partial masks, identifying contributions that a terminal score obscures. These two mechanisms make OPSD sensitive to how a trajectory should be corrected and where spatial credit should be assigned. Supervised mixing supports this self-teaching process by anchoring grounding and descriptive completeness to human annotations. Experiments demonstrate consistent improvements over the baseline across region understanding and localization. Pixel-OPSD thus advances pixel-level learning beyond scalar outcome feedback through trajectory-adaptive semantic correction and spatially localized credit, with all teaching mechanisms confined to training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.