DeSPO: Decoupling Rollout Budget from Set-Valuation Horizon in Policy Optimization
Abstract
Set-level objectives capture interactions among responses in reinforcement learning for large language models (LLMs). Yet scoring only the realized response set ties its valuation scope to the rollout budget. We introduce Decoupled Set Policy Optimization (DeSPO), separating the real rollout budget G from an independently chosen set-valuation horizon K. DeSPO optimizes probability weights over the generated responses using task rewards and the expected utility of K draws with replacement from the same fixed pool. The resulting target guides policy updates without additional text generation. For a determinantal utility, we derive an exact finite-support expansion showing that changing the horizon reweights interactions of different orders and is not generally equivalent to scaling a diversity bonus. Experiments across tasks, budgets, and model scales show that varying K at fixed G shifts learned quality–spread operating points, with effects that persist under fresh sampling. In a 3B setting, DeSPO achieves higher observed Mean RM and semantic spread than DQO with half as many real training rollouts. These results identify valuation horizon as a control axis distinct from rollout budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.