Resolvable-Effect Accounting for Tool-Calling Agents: Reward Bias and Evaluation Resolution
Abstract
Small gains from training tool-calling language models are difficult to assess when the training signal is affected by filtering and evaluation outcomes vary across runs. We introduce Resolvable-Effect Accounting (REA), a framework for examining reward logging and evaluation resolution, with candidate coverage and gradient statistics as supporting diagnostics. We show that, under group filtering in Group Relative Policy Optimization (GRPO), the retained-group reward is a conditional mean that need not move with the underlying success rate; logging the unfiltered mean avoids this selection effect for a fixed prompt distribution. In twelve-run HotpotQA evaluation panels, an estimated 69% of single-run variance is associated with run-level variation under the tested seed protocol, and same-seed pairing reduces comparison uncertainty when outcomes covary across conditions. In a Qwen3.5-2B case study, the observed checkpoint differences after 240 GRPO updates and 20 self-distillation updates are and percentage points, respectively, with 95% intervals that include zero. These results motivate explicit reporting of reward-selection rules and evaluation uncertainty. REA quantities can inform evaluation planning before optimizer updates, while their prospective value for training decisions remains to be assessed across matched configurations and additional settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.