acceptodds
Under review as a conference paper at ICLR 2027

Bridging Likelihood-Ratio Offline RL and Robust RL: Tractability and Error Bounds

Abstract

Likelihood-based offline reinforcement learning optimizes against models that fit the data, but its statistical appeal comes with a computational obstacle. A single likelihood constraint couples the entire model, whereas row-wise robust dynamic programming minimizes transition rows independently. We develop a tractable bridge between these formulations and quantify the price of using it. For tabular product transition models, planning over the exact state–action rectangular hull reduces to scalar reverse-KL problems and value iteration. Our central contribution is a two-sided, occupancy-sensitive analysis of the resulting approximation: a convex repair constructs a likelihood-feasible adversary, and a complementary convex program lower-bounds the loss incurred by independent row minimization. The error depends on incompatibility along visited rows, not likelihood-budget overspending alone. We characterize exactness for arbitrary initialization and distinguish fixed-policy approximation from coupled policy optimality. Combining this analysis with likelihood calibration yields offline performance guarantees that separate statistical uncertainty, relaxation error, and numerical error. The construction extends to valid shared linear-mixture models. Finite-model experiments support the bridge through faster matched-objective rectangular planning, occupancy-sensitive gap bounds, and common-parameter exactness checks. The bridge therefore makes a principled approximation to likelihood-based offline planning implementable and explains the additional pessimism introduced for computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.