Consensus-Breaking Re-rollout: Recovering Missing Action Contrasts for Long-Horizon Agent Reinforcement Learning
Abstract
Reinforcement learning for long-horizon language agents relies on limited rollout budgets, yet group-relative objectives obtain no learning signal from groups in which every rollout fails. Existing exploration strategies decide where to spend extra samples based on reward diversity or policy uncertainty. This overlooks a structural signal inside failed groups. We identify this signal as failure-conditioned consensus, where failed rollouts repeatedly select the same action at a shared decision point. Such agreement does not show that the action is correct. It marks a decision whose alternatives the current budget has never tested. To exploit this signal, we propose Consensus-Breaking Re-rollout (CBR), a budget-aware exploration framework that turns failed consensus into missing action contrasts. CBR operates through two complementary mechanisms. First, it locates consensus anchors within all-failure groups and re-rolls from these anchors while excluding the consensus action, so that alternatives are sampled within the support of the current policy. Second, it merges the recovered outcomes into the original group for two-level group-relative credit assignment. CBR requires neither process supervision nor a learned critic. In a fixed-policy diagnostic with two additional requests per all-failure group, CBR recovers a successful trajectory in of groups, compared with for ordinary retries and for entropy-guided branching. Across ALFWorld and WebShop with 1.5B- and 7B-parameter agents, CBR improves its base learner, GiGPO, by – percentage points at a matched generation budget. It also remains competitive with recent credit-assignment methods. These results validate failed consensus as an effective signal for allocating exploration in long-horizon agent learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.