acceptodds
Under review as a conference paper at ICLR 2027

PROSE: Enhancing Sample Efficiency in Code Generation via Plan-Aware Rollout Reorganization

Abstract

Group Relative Policy Optimization (GRPO) and related execution-based RL methods for code generation subtract a flat group-relative baseline from every rollout in a group; this baseline is variance-suboptimal as soon as rewards carry latent-plan structure. Programs instantiating the same latent plan tend to succeed or fail together, so execution rewards follow a one-way random-effects mixture whose within-plan intraclass correlation (ICC) is ANOVA-measurable ( on APPS+, on GSM8K). GRPO's i.i.d. rollouts remain marginally exchangeable, so the deficiency lies in the estimator and not in the sampling, but under this mixture the efficient subtractive baseline is plan-conditioned rather than flat. We propose PROSE, which couples explicit plan clustering, posterior-UCB budget reallocation across plans, and an allocation-aware training objective through a single shared plan posterior. Its hierarchical advantage is the uniformly minimum-variance unbiased (UMVU) estimator within the symmetric subtractive class, with a variance-reduction fraction set by . We test this prediction directly: at the training operating point (, ) the theory predicts an advantage-variance ratio of against the flat estimator recomputed on the same batch; the measured value on stored rollouts is . On LiveCodeBench pass@1, PROSE at () exceeds GRPO+CoT at () — strict superiority at half the execution budget — and PROSE at () exceeds GRPO+CoT at (); against vanilla GRPO, and in execution budget only, PROSE at matches GRPO at under a pre-specified equivalence margin of pass@1. On mathematical reasoning, where the ICC is lower, the gain is correspondingly smaller, as the diagnosis predicts.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.