ARBOR: Branch-Structured Credit Assignment for Reinforcement Learning with Verifiable Rewards
Abstract
Reinforcement learning with verifiable rewards has become the dominant recipe for eliciting long-form reasoning from language models, yet the reward it optimises arrives only after the final token has been emitted. Group-relative methods sidestep the cost of a learned critic by comparing whole trajectories sampled from the same prompt, but in doing so they collapse an entire reasoning trace onto a single scalar: every token in a correct rollout is reinforced, including the detours, and every token in an incorrect rollout is suppressed, including the steps that were sound. We show that this coarse credit is not an inevitable price of critic-free training but an artefact of how the group is sampled. \method replaces the flat group of independent rollouts with a rollout forest of the same token budget, forking continuations at positions where the policy is most uncertain, and reads a segment-level advantage directly off the branch structure as the difference between the empirical success rates of a node and its parent. The estimator requires no value network, no additional sampling, and no auxiliary reward model; it reduces to the group-relative advantage exactly when the forest degenerates to a flat group. We prove that the resulting gradient estimator is unbiased under a branch-mass reweighting of the leaves, that its variance is lower than the group-relative estimator's by an amount the branch structure determines, and that a depth-aware shrinkage attains the optimal bias–variance trade-off at finite branch counts. Across mathematical reasoning, code, and logic benchmarks and open-weight backbones from several families, \method improves over strong group-relative and critic-based baselines at matched sampling cost, sustains policy entropy far longer, and—unlike the baselines we compare against—does not narrow the solution coverage of the base model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.