When Right Meets Wrong: Bilateral Context Conditioning with Reward-Confidence Correction for GRPO
Abstract
Group Relative Policy Optimization (GRPO) has emerged as an effective method for training reasoning models. Although GRPO computes advantages from the group mean, it treats each output as an independent sample during optimization and leaves unused the contrast between correct and incorrect solutions within the same group, where successful reasoning traces could be compared explicitly against failed ones. To exploit this contrast, we present a contrastive reformulation of GRPO, showing that the GRPO objective implicitly maximizes the margin between the policy ratios of correct and incorrect samples. Building on this insight, we propose Bilateral Context Conditioning (BICC), a mechanism that allows the model to cross-reference successful and failed reasoning traces during the optimization, enabling a direct information flow across samples. We further introduce Reward-Confidence Correction (RCC) to stabilize training by dynamically adjusting the advantage baseline in GRPO using reward-confidence covariance derived from the first-order approximation of the variance-minimizing estimator. Both mechanisms operate on the sampled group alone and apply to GRPO and its variants. Experiments with Qwen3-4B and Phi-4-mini on Math500, AMC 2023, and AIME 2024/2025 show that BICCimproves GRPO and five of its variants in 54 of 56 comparisons, by 0.3 to 1.9 points with larger gains on the weaker base model, while RCC brings gradient variance 31 to 37% below GRPO. The gains persist at 14B scale and on Llama-3.1-8B, and training on mathematics transfers zero-shot to logic and code. The implementation is available at [BiCC Codebase](https://anonymous.4open.science/r/BiCC-77C9)
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.