acceptodds
Under review as a conference paper at ICLR 2027

From Rollout Scores to Policy Updates: Estimator Completion for Agentic RL

Abstract

Agentic reinforcement learning often scores groups of policy-generated rollouts before selecting which ones enter an update. Selection compresses graded priorities into a mask, leaving open how the discarded score mass should affect the estimator. We formulate this missing step as Estimator Completion (EC). Given a scorer and an optimizer's native aggregation on selected groups, conserving the selected conditional, the scorer's tail conditional, and their partition mass uniquely determines a completed group measure. It is a KL projection of the score distribution, differs from hard filtering by exactly the omitted mass, and yields an explicit head-recalibration and tail-restoration correction. Reward-dispersion filtering provides a concrete test: logged batches retain ordered scores and reward-active rejected groups, and we carry EC's score-conditioned coordinates into PPO, GRPO, DAPO, and DrGRPO policy reductions. Across four tasks, all four Qwen2.5-3B EC variants improve the four-task average of peak-validation means over RAGEN-2; EC-PPO does so across all six tested models. This gives graded rollout selection a principled estimator interface beyond its keep/drop rule.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.