acceptodds
Under review as a conference paper at ICLR 2027

Selection Is Not Evaluation: Maximization Bias in Scoring Runtime Interventions for Language Agents

Abstract

Runtime interventions for language agents, such as a warning, a new plan, or a rollback, are often judged by branching from a shared prefix and crediting the intervention whose rollout succeeds. Because rollouts are stochastic, this credits the best of several noisy outcomes, the same maximization bias that motivated Double Q-learning. We propose a repeated-branch protocol that gives every action, including no intervention, two independent rollouts from the same state, selects the action on one and scores it on the other, and fixes all controller predictions before any outcome is observed. On held-out ALFWorld tasks with a Qwen3-8B actor, scoring on the same rollout overstates the value of intervening by 15.4 percentage points. With this bias removed, a controller trained to predict each intervention's benefit over continuing gains less than two points over never intervening. That difference is not significant, and the controller does no better than one that predicts per-action outcomes. On ScienceWorld the gap is smaller, and no controller improves on continuing. A rescue seen on a single rollout is therefore weak evidence that intervention helps. Evaluations should include continuing as an action and score each choice on a rollout that was not used to make it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.