acceptodds
Under review as a conference paper at ICLR 2027

TACE: Two-Stage Advantage and Credit Estimation for Multimodal Reinforcement Learning with Verifiable Rewards

Abstract

We revisit advantage estimation in Group Relative Policy Optimization (GRPO) and observe that any token-level advantage admits a two-stage factorization: a sample-level baselining step that converts scalar verifier rewards into a comparable sequence-level signal, and a token-level credit assignment step that distributes that signal over the trajectory. Existing GRPO variants modify only one of the two stages, and we show that each leaves a structural failure mode in the other untouched: linear standardization preserves the shape of the empirical reward distribution, while teacher-driven distribution matching encodes privileged context into the gradient direction and is provably ill-posed under information asymmetry. We propose TACE (Two-stage Advantage and Credit Estimation), a single estimator that fixes both stages. The first stage replaces linear normalization with the unique 1D optimal-transport map onto a standard normal, yielding inter-task gradient equity, robustness to heavy tails and bimodality, and strict sign preservation. The second stage uses an evidence ratio between the model under privileged and non-privileged contexts as a stop-gradient, sign-anchored magnitude multiplier, ensuring that the verifier retains exclusive control of the gradient direction. Composing the two stages costs only one sort and one extra forward pass per group beyond standard GRPO. On Qwen3-VL-Instruct-8B, TACE outperforms the base model by 4.69 points and the strongest GRPO variant by 2.32 points averaged over MMMU, MathVista, MathVision, ZeroBench, and WeMath, while sustaining improvement well past the point where teacher-driven self-distillation degrades.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.