acceptodds
Under review as a conference paper at ICLR 2027

CoVER: Reinforcement Learning with Code-Verified Process Rewards in Stateful Code Execution Environments

Abstract

Large language models increasingly rely on external code execution for mathematical and code-intensive reasoning. However, most RL post-training assigns credit only through final-answer correctness, while prior process-supervision methods are largely designed for static traces and often require costly step labels or learned evaluators. We introduce **CoVER** (**Co**de-**Ve**rified Process **R**ewards), a reinforcement learning framework that turns stateful code execution into online process supervision. CoVER trains LLMs through iterative execution and self-correction in a persistent sandbox, using lightweight process rewards derived from execution-grounded sandbox signals. To optimize mixed agent–environment trajectories, we propose PR-GRPO, a process-reward variant of GRPO that applies masked log-probabilities and masked KL regularization only to policy-generated tokens and assigns credit with depth-wise step-level advantages. We construct CoVERBench by standardizing public reasoning sources and solver-verified tasks into a unified benchmark spanning mathematical reasoning, logical deduction, and constraint satisfaction. On CoVERBench, CoVER improves Qwen3-8B to 93.7% held-out validation accuracy, versus 82.1% for outcome-only GRPO under a matched protocol, with non-overlapping bootstrap confidence intervals. CoVER also transfers to unseen math and code benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.