acceptodds
Under review as a conference paper at ICLR 2027

Recovering Verifiable Process Credit in Reinforcement Learning for Code Generation

Abstract

Group-based reinforcement learning with verifiable rewards (RLVR) and its recent advances train a code model by examining the outcomes of sampled rollout programs against a suite of test cases. It thus suffers from the sparse-reward issue: a reward is credited at the level of the whole program. In contrast to this coarse-grained perspective, we view a program as a composition of syntax-tree units and find that even a failing program often contains units that already make tests pass, and that its failure can be traced to a few specific units. Based on the findings, we propose ASTrace, an abstract syntax tree (AST) reward attribution method that assigns process credit to program units from code execution. ASTrace parses each sampled program into units, namely its functions, methods, and top-level blocks, and credits them through two components: (i) *progress credit*, which, for all-fail rollouts, credits each unit with its marginal gain in passing tests; and (ii) *failure localization*, which, in partial-fail rollouts, penalizes the unit of a failing rollout where a missed test raises its error. We compare ASTrace against GRPO and its most recent variants across six code-generation benchmarks and two backbones. ASTrace consistently improves pass@1 over GRPO on both common coding benchmarks (e.g. 2.2% on HumanEval with Qwen3-4B) and more challenging coding contests (e.g. 2.6% on OJBench with Qwen3-4B). Our analysis shows that ASTrace turns near-miss programs into solutions and gains the most where failing programs are partially correct.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.