acceptodds
Under review as a conference paper at ICLR 2027

Objective-Aware Gaussian Sampling for Zero-Variance Prompts in LLM Reinforcement Learning

Abstract

Group Relative Policy Optimization (GRPO) has emerged as an efficient reinforcement learning paradigm for fine-tuning large language models, yet its effectiveness can be hindered by zero-variance prompts—cases where all sampled responses yield identical rewards, producing zero group-relative advantages while still incurring generation costs. We introduce Objective-Aware Gaussian Sampling (OAGS), a prompt-level curriculum framework that dynamically reallocates the rollout budget toward prompts in an evolving target pass-rate region. OAGS estimates prompt-specific pass rates from observed outcome rewards, constructs Gaussian candidate proposals centered on the current curriculum target, and adaptively adjusts the curriculum via a progress-controlled pacing mechanism. By concentrating training on informative prompts while preserving controlled exploration of harder and easier regions, OAGS prioritizes non-zero-variance response groups within a fixed candidate budget. Our approach retains the GRPO advantage estimator and clipped surrogate while adapting prompt sampling and post-rollout group selection. Experimental results show improved macro-average performance across six mathematical reasoning benchmarks compared to standard GRPO and the evaluated data-selection baselines, with fewer training rollouts than Dynamic Sampling and GRESO at matched policy-update counts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.