Dirichlet-GRPO: Posterior-Calibrated Advantage Transport for Multimodal RLVR
Abstract
Multimodal reinforcement learning with verifiable rewards converts small groups of discrete outcomes into policy-update weights, introducing error through uncertain outcome frequencies and approximate representatives of the resulting probability regions. To address both sources, we propose \method, which combines Jeffreys–Dirichlet posterior mean probabilities with exact Gaussian interval means so that probabilities determine each ordered region's area and its mean determines the update weight. We establish the optimal representative for a fixed partition, characterize the bias–variance trade-off of probability smoothing, and bound how probability error affects the initial policy update, while exact controlled experiments isolate the two corrections and identify conditions where smoothing is unfavorable. Across the 2B, 4B, and 8B scales of Qwen3-VL, our \method improves upon GRPO by 2.77, 1.37, and 0.76 points on average. On 4B scale, it surpasses both GRPO and GRPO across all 11 evaluated benchmarks, notably boosting WeMath, MMMU, and MMStar. To disentangle these gains, exact Bernoulli ablations isolate the impact of posterior calibration and Gaussian interval integration, showing up to a 37.0% reduction in initial gradient MSE over tied ranks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.