acceptodds
Under review as a conference paper at ICLR 2027

Down-Weighting Uncertain Supervision: Frozen Checkpoint-Label Compatibility Weighting for GRPO

Abstract

Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective paradigm for fine-tuning vision-language models on classification tasks, yet the standard objective does not explicitly distinguish the reliability of the observed supervision associated with different training samples. We study SCS-GRPO, which assigns each sample a frozen checkpoint-observed label-compatibility score and applies it as a post-normalization multiplier on the GRPO advantage. We analyze why applying a prompt-level scalar before within-group normalization can largely remove its intended effect, and give a conditional gradient signal-to-noise analysis under explicit assumptions about corrupted supervision and the compatibility ranking. Experiments on multiple public benchmarks show consistent improvements in the evaluated classification and noise settings. An anonymous implementation is available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.