acceptodds
Under review as a conference paper at ICLR 2027

Offline Policy Learning from Outcome-Only Feedback: Identification and the Price of Logging Memory

Abstract

We study offline reinforcement learning when each logged episode records the full trajectory but only one noisy scalar outcome, with no per-step rewards. Under history-dependent logging, state–action coverage alone may not determine whether a target policy can be evaluated, because the data reveal rewards only through sums along logged trajectories. With known dynamics, we show that a policy value is identified when its occupancy lies in the span of logged trajectory features; for interior tabular rewards, this condition is also necessary, and when point identification fails the value admits a sharp identified interval. The coefficient is the exact additive change-of-measure constant and quantifies finite-sample difficulty. History-dependent logging separates this trajectory geometry from ordinary state–action coverage. Two loggers can have identical one-step occupancies while combining those visits into different trajectories, changing both identification and . We quantify this effect through an order-necessary distortion factor and a target-specific refinement . With known dynamics, we propose a projected, box-constrained pessimistic learner whose expected regret depends on of an optimal policy without requiring a global minimum-eigenvalue threshold; block splitting gives a corresponding high-probability guarantee. On a temporal direct-sum family, exactly, with complementary Gaussian and missing-exposure lower-bound mechanisms. For general even horizons, a parallel-state construction gives the minimax rate for the stated high-dimensional tabular class with under the non-clipping conditions. Matched-marginal experiments on compiler-derived additive geometry and at the Markov logging boundary confirm that changing cross-time dependence changes outcome geometry and learning difficulty.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.