acceptodds
Under review as a conference paper at ICLR 2027

Semiparametrically Efficient Off-Policy Evaluation and Learning in Constrained Markov Decision Processes

Abstract

We study off-policy evaluation (OPE) in constrained Markov decision processes (CMDPs): estimating the reward value and the constraint cost values of a new policy from data collected by another. Before a policy is deployed as safe, its feasibility must be certified from confidence intervals, not point estimates. Such a certificate concerns a combination of the estimates—typically the Lagrangian —so its validity depends on how the reward and cost errors covary. Existing methods fall short in two ways. Concatenating the step-wise doubly robust estimator across objectives inherits its exponential variance in the horizon. Applying a marginally efficient estimator to each objective separately breaks the curse of horizon, but treating the objectives as independent for inference discards the covariance induced by the marginal density ratio they share. We instead treat the reward and cost values as a single vector target, derive its efficient influence function, and propose the Constrained Doubly Robust (CDR) estimator. CDR is unbiased whenever either the density ratio or the -functions are correctly specified, attains the joint semiparametric efficiency bound under a product-rate condition, and reports the full reward–cost covariance from one cross-fitted pass. The sign of that covariance governs the failure this repairs: when positive, independent treatment is merely conservative; when negative, it is anti-conservative—the Lagrangian intervals under-cover and an unsafe policy can be certified. Experiments on tabular and continuous-state CMDPs, and on a constrained Taxi benchmark at horizon , confirm the theory, including the negative-covariance regime and a diagnostic in which violating the product-rate condition collapses coverage.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.