CodeVerifier: Unified Generative Verification across Code Generation and Development
Abstract
Code generation for competitive programming and repository level software engineering can be viewed as a sequence of decisions over code states such as complete programs ongoing edits and repository patches. Reinforcement learning with verifiable rewards can improve these decisions beyond finite demonstration data provided that the policy receives a reliable reward before selecting its next action. This condition is difficult to meet at scale since short test prefixes leave many candidates tied deeper execution pushes latency into a long tail and intermediate edits often lack a testable entry point while querying a larger general purpose model at every step is too costly. We turn this missing feedback into a reusable signal with CodeVerifier, a generative verifier that judges code states before final testing and outputs a rationale repair relevant code regions and a verdict. Supervision comes from two kinds of history where execution outcomes indicate whether a state succeeds and successful repairs indicate where a failed state must change. A two stage procedure first projects this evidence into structured targets and then optimizes verdicts and regions with separate field wise signals. CodeVerifier outperforms size matched code judging baselines across competitive programming and repository code by points on average. On LiveCodeBench residual errors, one 9B judgment matches the mean detection rate of 8 additional tests; on the hardest TACO slice, matched-recall execution is slower even with idealized 64-way parallelism. Used as a reward it improves candidate selection distinguishes successful from unsuccessful edits and raises final task success which shows that structured early judgments can fill the feedback vacuum and replace many repeated executions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.