acceptodds
Under review as a conference paper at ICLR 2027

Bellman Recursion Diffusion for Off-Policy Evaluation

Abstract

Off-policy evaluation (OPE) estimates the value of an evaluation policy from data collected by a different behavior policy. A generative model of the policy's discounted future answers value queries without sequential rollouts, but training one by bootstrapping on its own predictions amplifies fitting error by up to the effective horizon . We introduce Bellman Recursion Diffusion (\method), which trains a conditional diffusion model on the discounted distribution of future transitions of a frozen one-step dynamics model. Each training target lets the frozen model simulate the first transitions and bootstraps only the remaining mass , which reduces the amplification to . We prove that every probability-valued Bellman target buys contraction at the same rate per model call, so depth decides only where the horizon is paid for; that the Bellman structure holds at every noise level, so denoising score matching fits it; and that the estimator obeys an OPE bound separating generative from dynamics error. On five control benchmarks, deeper blocks bring the generator closer to the frozen model, and on Walker2d one-step bootstrapping remains far from it even after fifty times more sweeps. A released one-step occupancy model, TD-flow, has higher log-RMSE than \method on all five tasks and lower rank correlation on four. \method is more accurate on both metrics than model-based rollout on all four horizon-matched tasks, and than STITCH-OPE on three of them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.