HAVER-GRPO: Hazard-Aware Rollout Allocation for Efficient RLVR
Abstract
Reinforcement learning with verifiable rewards can waste substantial rollout compute on prompt groups whose responses are all incorrect and yield no group-relative learning signal. We introduce , a sequential allocation strategy that decides whether an all-wrong group merits another rollout. HAVER estimates verifier-success hazard and response-token cost, then allocates fresh rollouts by marginal recovery yield under a shared budget price. The first verified response is inserted into a fixed-size GRPO group, while every attempted token is charged to the compute budget. We analyze recovery probability, expected cost, and sensitivity to estimation error. Across five mathematical reasoning benchmarks, our main experiments achieve mean accuracies of and for the 1.5B and 7B models, respectively. Tokens-to-target decrease by and relative to uniform GRPO, demonstrating improved training efficiency at the same accuracy targets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.