acceptodds
Under review as a conference paper at ICLR 2027

The Cost of Evaluating Policy Updates under Censored Feedback

Abstract

A policy can appear more accurate simply because its failures are less likely to be observed. We study how to decide whether a proposed policy actually improves expected reward when some rewards are missing. An audit reveals one missing reward, but the corresponding outcome must first be generated. Evaluation therefore has two costs: producing outcomes and inspecting them. For finitely many actions and rewards in [0,1], we characterize the smallest worst-case estimation error under both budgets. The optimal allocation gives more audits to influential actions, subject to how often their outcomes are available. For actions with too few observations, it reduces reliance on the sampled rewards and includes the remaining uncertainty in the error bound. We also account for estimating missingness and for evaluating a sequence of proposed updates. Experiments in a controlled environment and on fixed Qwen-generated GSM8K responses show when additional evidence resolves more decisions, and why that need not improve the final policy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.