acceptodds
Under review as a conference paper at ICLR 2027

CARVE: Credit Assignment via Routed Verification of State-Conditioned Behavioral Units

Abstract

Reinforcement learning with verifiable rewards (RLVR) provides outcome-level supervision for reasoning models, but many reasoning and decision-making tasks expose intermediate verification signals during generation. Effectively exploiting these signals requires identifying appropriate intermediate credit units, matching behaviors with informative verification, and translating heterogeneous feedback into localized policy updates. We propose CARVE, a framework for credit assignment via routed verification of state-conditioned behavioral units. CARVE organizes trajectory tokens into structurally segmented and behaviorally annotated units, each associated with its preceding reasoning, execution, or interaction state. It then applies behavior-specific verification to assess local validity and continuation-based verification to estimate how selected units change the state-dependent downstream utility. Both evidence sources are converted into separate behavior-level credits and localized to the corresponding tokens, while the final task reward remains the primary optimization signal. CARVE outperforms the strongest evaluated baselines across two model families using only 400 curated training problems per domain. It achieves gains of up to 7.00 percentage points in mathematical reasoning, 3.66 in code generation and 2.50 in multi-turn tool use. These results demonstrate its effectiveness on challenging tasks with multi-step dependencies and state-dependent decisions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.