Safe to Switch: Target-Covered Off-Policy Evaluation via Policy Replacement
Abstract
Offline reinforcement learning enables policy learning from existing data when further interaction is costly or unsafe. However, target occupancy coverage and linear realizability of the target's action values alone do not guarantee sample-efficient off-policy evaluation. We show that target coverage suffices under replacement-policy linear realizability, using safe partial corrections to logged returns. Replacement policies retain a drawn logger action or replace it with a fresh target draw, with arbitrary probabilities across stages, states and drawn actions. At every stage, every replacement policy's action values must be linear on all state–action pairs in known features, with uniform feature and parameter norm bounds. For a fixed target, our estimator achieves error at most with probability at least using iid complete trajectories from one unknown fixed Markov logger, without logger probabilities or model oracles. Here bounds target-to-logger occupancy ratios, is the horizon and the feature dimension. The sample bound has logarithmic norm dependence and no state or action cardinality factor. Separate families establish strict representation and coverage relaxations: replacement-policy values admit a two-dimensional feature map while all-policy dimension is unbounded, and finite target coverage coexists with infinite mixture coverage. The finite-sample guarantee and both separations are formalized in Lean 4.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.