Higher-Order Exploration in RLOO with Pairwise-Independent Rollouts
Abstract
Sparse rewards can leave an entire rollout group without a learning signal. We introduce HO-RLOO, an autoregressive sampler that induces higher-order dependence while preserving the independent law of every trajectory pair and therefore the expected raw RLOO gradient under ideal continuous randomness. A calibrated mixture of distinct and shared probability strata changes interactions among three and four trajectories without changing pairwise laws. In an aligned sparse-reward model, HO-RLOO reduces gradient variance and achieves maximal group coverage among pairwise-independent samplers. On mathematics and code, HO-RLOO improves domain-balanced pass@1 over iid-RLOO by 2.51 and 2.78 points at two Qwen3 model scales under a common training-token cap. Frozen-policy diagnostics show increased mixed-reward groups and rollout diversity on the evaluated prompt panel. These results demonstrate that preserving pairwise independence provides a principled way to shape higher-order exploration while retaining the original RLOO objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.