acceptodds
Under review as a conference paper at ICLR 2027

Model Selection for Off-Policy Evaluation via Pairwise Bellman Validation

Abstract

Model selection is a central challenge in off-policy evaluation (OPE) because fitted -functions and occupancy ratios lack supervised validation losses. Standard approaches based on Bellman-residual estimation or minimax criteria often require auxiliary regressions or critic classes that must themselves be learned and tuned. We introduce pairwise Bellman validation (PBV), a tuning-free selection procedure among a finite collection of fitted -functions and occupancy-ratio estimators that uses only held-out transitions and requires no auxiliary validation models. The construction is motivated by an oracle critic given by a candidate's estimation error, which, when used to test the Bellman residual, controls the error of interest. Because this oracle critic is unknown, we approximate it using pairwise differences between candidates. Under misspecification, these pairwise approximations are imperfect, so we add penalties that account for the resulting critic-approximation error. For -function selection, we introduce target-aligned PBV, which uses marginalized importance weighting to evaluate the pairwise tests under the discounted target-occupancy distribution. When occupancy-ratio estimation is undesirable, we propose an unweighted variant closely related to LSTD-Tournament, but it requires an additional conditioning assumption linking offline Bellman tests to -function error. For occupancy-ratio selection, we introduce adjoint PBV, which applies the same pairwise construction to the adjoint Bellman equation. For all three procedures, we establish population and finite-sample oracle inequalities relative to the best candidates in possibly misspecified classes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.