acceptodds
Under review as a conference paper at ICLR 2027

Resolvable-Effect Accounting for Tool-Calling Agents: Reward Bias and Evaluation Resolution

Abstract

Small gains from training tool-calling language models are difficult to assess when the training signal is affected by filtering and evaluation outcomes vary across runs. We introduce Resolvable-Effect Accounting (REA), a framework for examining reward logging and evaluation resolution, with candidate coverage and gradient statistics as supporting diagnostics. We show that, under group filtering in Group Relative Policy Optimization (GRPO), the retained-group reward is a conditional mean that need not move with the underlying success rate; logging the unfiltered mean avoids this selection effect for a fixed prompt distribution. In twelve-run HotpotQA evaluation panels, an estimated 69% of single-run variance is associated with run-level variation under the tested seed protocol, and same-seed pairing reduces comparison uncertainty when outcomes covary across conditions. In a Qwen3.5-2B case study, the observed checkpoint differences after 240 GRPO updates and 20 self-distillation updates are and percentage points, respectively, with 95% intervals that include zero. These results motivate explicit reporting of reward-selection rules and evaluation uncertainty. REA quantities can inform evaluation planning before optimizer updates, while their prospective value for training decisions remains to be assessed across matched configurations and additional settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.