acceptodds
Under review as a conference paper at ICLR 2027

PREDICTIVE, NOT REACTIVE: TASK FILTERING FOR COMPUTE-EFFICIENT GRPO TRAINING OF CODE AGENTS

Abstract

Group Relative Policy Optimization (GRPO) learns by comparing several rollouts on the same task. If every rollout receives the same reward, their relative advantages are zero. The task then provides no policy gradient, even though all of its rollouts have already been generated. This is common in repository-level code repair for two reasons. A binary reward treats a partial repair as a complete failure: a patch that fixes some failing tests receives the same reward as one that fixes none. In addition, tasks that are too difficult for the current policy often give the same reward for every rollout. In both cases, GRPO spends computation on a group from which it cannot learn. We propose PRESAGE (PREdictive SAmpling with Graded Evaluation), which combines task filtering with a per-test reward. Before generating rollouts, a pass-rate predictor reads the task description and estimates the score that the current policy would obtain. Tasks with predicted scores below a threshold are skipped, so no rollout budget is spent on them. For the remaining tasks, the reward gives credit for the fraction of failing tests repaired and subtracts a penalty for breaking tests that previously passed. This allows different partial repairs to receive different scores. A task can therefore provide a learning signal even when no rollout solves it completely, as long as the rollouts make different amounts of progress. These per-test scores are also used to train the predictor. Because the policy improves during training, PRESAGE recalibrates the predictor and checks the full task pool again in every cycle. A task skipped earlier can return once the policy becomes strong enough to make progress on it. This keeps later batches from filling up with tasks that produce no gradient. With the same task pool, rollout budget, and reward function, predictive filtering increases the fraction of sampled groups that provide gradients from to . The same budget supports 40 policy updates with PRESAGE, compared with 7 under dynamic sampling. Compared with uniform GRPO, PRESAGE improves the fail-to-pass rate by percentage points and the resolved rate by points. Compared with the supervised fine-tuned starting model, the improvements are and points, respectively.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.