acceptodds
Under review as a conference paper at ICLR 2027

HAVER-GRPO: Hazard-Aware Rollout Allocation for Efficient RLVR

Abstract

Reinforcement learning with verifiable rewards can waste substantial rollout compute on prompt groups whose responses are all incorrect and yield no group-relative learning signal. We introduce , a sequential allocation strategy that decides whether an all-wrong group merits another rollout. HAVER estimates verifier-success hazard and response-token cost, then allocates fresh rollouts by marginal recovery yield under a shared budget price. The first verified response is inserted into a fixed-size GRPO group, while every attempted token is charged to the compute budget. We analyze recovery probability, expected cost, and sensitivity to estimation error. Across five mathematical reasoning benchmarks, our main experiments achieve mean accuracies of and for the 1.5B and 7B models, respectively. Tokens-to-target decrease by and relative to uniform GRPO, demonstrating improved training efficiency at the same accuracy targets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.