Confidence-Guided Policy Optimization: Thompson Sampling and UCB Trees for Group-Relative Code Refinement
Abstract
Group-relative policy optimization methods, such as GRPO, learn from the relative reward of trajectories sampled for the same prompt. When every trajectory in a group receives the same reward, the centered advantages vanish and the group contributes no policy-gradient signal, even though the environment was queried and feedback was obtained. We introduce Confidence-Guided Policy Refinement (CGPR), a training-time framework that uses this feedback to convert such uninformative groups into informative ones, rather than discarding them and resampling independently. CGPR builds a bounded refinement tree over the trajectories of an uninformative group and evaluates two budget-allocation strategies for expanding that tree: a UCB-style rule (CVPR) and a Thompson-Sampling rule (CGPR-TS) based on a Beta posterior over per-node success probability. A candidate replaces a non-anchor trajectory only when it increases the reward dispersion of the group, and a highest-reward anchor trajectory is never replaced. Refinement occurs only during training and adds no inference-time computation. All experiments use the same Qwen3-4B-Instruct backbone. Across HumanEval(+), MBPP(+), APPS, Codeforces, and LiveCodeBench, the two CGPR variants improve over vanilla GRPO on the more challenging APPS and Codeforces benchmarks, while their effects on HumanEval(+), MBPP(+), and LiveCodeBench depend on the metric and allocation strategy. Relative to GRPO, the largest pass@1 and pass@10 gains are both percentage points, achieved by CVPR on Codeforces. CGPR-TS obtains the strongest MBPP(+) pass@1/pass@10 and LiveCodeBench pass@1 among the fine-tuned methods, while CVPR obtains the strongest APPS and Codeforces results. DAPO obtains the strongest HumanEval(+) pass@1/pass@10 and LiveCodeBench pass@10 among the fine-tuned methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.