When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?
Abstract
Offline reinforcement learning is typically analyzed under process-level reward supervision, yet many sequential decision datasets record only trajectory-level outcomes, such as task success, a final score, or a human preference between two trajectories. We develop a statistical theory of offline policy optimization from such feedback. For the cumulative-reward objective with one scalar outcome label per trajectory, we propose OPAC, a pessimistic actor-critic algorithm that learns a latent per-step reward together with the critic. OPAC finds an -optimal policy from trajectories under single-policy concentrability , and we prove a matching lower bound, characterizing the statistical cost of replacing process-level rewards with one trajectory-level label. The approach extends to pairwise trajectory preferences with comparable guarantees. We then study nonlinear aggregations of per-step rewards that serve as both the supervision and the objective. Here efficiency is not guaranteed: all-success outcomes can require exponentially many trajectories even with deterministic transitions and constant coverage. For structured aggregations affine in the continuation value, which admit Bellman-style dynamic programming, we identify two coefficients that quantify the information lost in outcome aggregation and Bellman updates, and show that generalized OPAC is sample-efficient when both are polynomially bounded.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.