Adversarial Reinforcement Learning with Transition Predictions
Abstract
We study episodic adversarial Markov decision processes with unknown transition dynamics, where the learner is additionally given a potentially misspecified prediction of the transition kernel. We propose Transition-Prediction Follow-the-Regularized-Leader (TP-FTRL), an occupancy-based algorithm that combines empirical transition confidence sets with likelihood-calibrated confidence sets induced by the prediction. We establish simultaneous high-probability coverage of the true transition kernel without assuming prediction accuracy. We further prove a prediction-independent worst-case regret bound of , matching the state-of-the-art bound by Jin et al. (2020), and showing that arbitrarily inaccurate transition predictions do not degrade the worst-case regret guarantee. When the prediction is close to the empirical transition estimates, the likelihood-calibrated constraints tighten transition uncertainty and yield a sharper regret bound. These results show that TP-FTRL can exploit informative transition predictions while remaining robust to misspecification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.