acceptodds
Under review as a conference paper at ICLR 2027

Process Coherence Games for Code Alignment: Bootstrapping from Zero

Abstract

Large Language Models (LLMs) excel at code generation and complex multi-step reasoning, yet aligning them toward functional correctness remains a severe challenge. Standard alignment paradigms, such as Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), rely on syntactic matching against canonical reference traces, thereby penalizing valid but syntactically diverse reasoning paths. Conversely, outcome-based Reinforcement Learning (RL) suffers from extreme gradient variance and produces zero gradient signal in complex domains where the probability of an initial correct solution is near zero (). In this paper, we overcome these bottlenecks by formulating code generation alignment via *Pairwise Coherence*. We propose C2-GRPO and its probabilistic extension PC-GRPO, which bypass the need for exhaustive ground-truth oracles by optimizing for functional semantic consistency over dynamically generated unit tests. To solve the cold-start problem, we introduce two algorithmic techniques, *Soft Plausibility Weighting* and *Process Coherence*, which dynamically synthesize an implicit Process Reward Model (PRM) entirely through unsupervised group consensus, reducing the group size needed for a consensus signal on long-horizon tasks from exponential to linear under step-wise decomposition (via tree-based simulation or, for code generation, test-decomposed independent rollouts). We prove that standard GRPO's gradient signal vanishes as while PC-GRPO maintains advantage variance whenever the group's functional disagreement is non-degenerate (an explicit count-separation condition), and establish convergence and Rademacher generalization bounds for an idealized form of the coherence objective. Empirically, on Qwen2.5-Coder-7B-Instruct, PC-GRPO is the strongest trainable method across APPS Intro, APPS Interview, and LiveCodeBench release_v1. Even at 7B scale, APPS Interview retains a substantial cold-start regime; in these groups, PC-GRPO produces usable gradient variance in 91.0% of cases where binary rewards provide none. We further prove fundamental impossibility results for ungrounded consensus (the *Degenerate Consensus* collapse) and its partial converse (*grounded one-step improvement*), and integrate bounded Minimax Polynomial estimators (Mohri et al., 2026b) for finite-sample stability, and structurally bound reward hacking via the consensus formulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.