DSE-GRPO: Differential Semantic Entropy for Debiased Noisy-Rollout Policy Optimization
Abstract
Reinforcement learning (RL) post-training for vision–language models (VLMs) increasingly relies on noisy rollouts to encourage exploration and improve robustness. However, mixing rollouts generated from clean and perturbed images introduces a hidden off-policy mismatch: noise can change the semantic interpretation of the input, yielding biased and high-variance policy gradients when these trajectories are optimized together under GRPO. We propose DSE-GRPO, an uncertainty-aware variant of GRPO that explicitly measures and suppresses noise-induced semantic drift. Our key idea is Differential Semantic Entropy (DSE): we estimate semantic entropy over response clusters for clean and noisy rollout groups separately, and use their non-negative gap as a differential uncertainty signal. DSE-GRPO then adaptively down-weights the advantages of noisy rollouts with large DSE, filtering harmful off-policy updates while retaining beneficial exploratory trajectories. We provide theoretical analysis showing that DSE-based reweighting attenuates off-policy bias and stabilizes optimization dynamics, while slowing policy-entropy collapse to preserve exploration. Empirically, RL-tuning with Geometry3K and MMK12 shows that DSE-GRPO consistently outperforms vanilla GRPO and NoisyRollout baselines, improving the average accuracy of Qwen2.5-VL-7B and Qwen3-VL-8B by 8.0 and 1.9 percentage points, respectively, across five out-of-domain benchmarks, with markedly smoother training dynamics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.