CABRA: Cost-Aware Bayesian Rollout Allocation
Abstract
Group-relative reinforcement learning can spend much of its computation on response groups whose rewards barely differ. Discarding such groups after generation does not recover their sampling cost, and searching until a fixed number of useful groups is found can require many additional prompts. We introduce Cost-Aware Bayesian Rollout Allocation (CABRA), which fixes each prompt's number of responses before generation. From each prompt's past outcomes, it estimates, per expected token, the probability that a group will contain both correct and incorrect responses. A fixed number of top-ranked prompts receive complete groups of the full dense size. Skipped prompts with past outcomes then receive two-response refills if their expected token cost fits the remaining budget. All groups are generated once, without searching for replacements. On Qwen2.5-1.5B and GSM8K, CABRA uses 38.5% less adapter-backward work (LoRA backward-pass FLOPs) than dense DAPO across five controlled seeds, with 623.2 versus 621.4 mean correct answers out of 1,319. Without refills, work falls by 52.2% with 622.2 mean correct answers, so most of the saving comes from limiting the number of complete groups, and refills spend part of it. Generated tokens follow the same ordering. The accuracy differences are not statistically significant. We also compare CABRA with five external allocators in an eight-prompt protocol.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.