Choosing What to Remember: Robust Offline Reinforcement Learning with History Representations
Abstract
Robust offline reinforcement learning uses recorded data alone to learn policies that perform reliably under uncertainty. With delays or partial observations, policies often use history representations that map observed histories to represented states. The joint distribution of reward and next represented state can then depend on the policy, even at a fixed represented state and action in an unchanged environment. We develop a framework for conditional operator robustness under history representations and propose Conditional Operator Robust Fitted Q-Iteration (COR-FQI). Its robust Bellman recursion chooses actions that maximize the worst-case expected sum of immediate reward and continuation value over uncertainty sets. These sets are constructed to account for differences in conditional Bellman expectations among histories sharing a represented state, in addition to finite-sample estimation error and specified environmental changes. Under stated assumptions, we establish lower confidence bounds on the learned policies' worst-case expected returns without requiring the represented process to be Markov. The guarantees cover tabular settings and function approximation, with additional coverage conditions and uniform error bounds for the latter. These bounds hold simultaneously across a finite family of representations and provide a performance guarantee for choosing the representation–policy pair with the largest lower confidence bound. Tabular experiments show informative bounds and selection of longer memory with more data, while implementations based on function approximation show empirical robustness to perturbations, delays, and distribution shift.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.