AUDITED PARTIAL IDENTIFICATION FOR OFF-POLICY EVALUATION UNDER NONIGNORABLE FEEDBACK
Abstract
When reward recording depends on the unseen reward itself, even unlimited logged data may not identify policy values. We study off-policy evaluation in this setting and ask which missing rewards are most informative to recover through targeted audits. Under observed transitions, target coverage, and a shared rectangular reward model, we characterize the sharp joint set of policy values and quantify how common uncertainty cancels in policy comparisons. At the population level, identifying a cell’s reward mean reduces comparison width by exactly the product of the policies’ absolute visitation difference and the cell’s reward-mean ambiguity, independently of the recovered mean. Minimizing the remaining width therefore reduces to top- audit selection for equal costs or a knapsack problem for unequal costs. For a fixed observation design, half this width lower-bounds the worst-case expected absolute estimation error. An independent design–confirmation procedure provides asymptotically valid outer confidence intervals, including when policy visitation differences are zero or near zero. Compared with separate marginal bounds, shared-law comparison yields 19.3-fold narrower population bounds in a controlled MDP and 4.4-fold narrower plug-in bounds in a semi-synthetic KuaiRec study. Two real follow-up analyses further demonstrate the importance of outcome stability and accounting for incomplete reward recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.