Almost Exogenous: Regret Guarantees for Reinforcement Learning under Bounded Action Influence
Abstract
Exogenous Markov Decision Processes (Exo-MDPs) model settings where uncertainty is driven by external factors independent of the learner's actions. In many real systems, this separation is only approximate: actions weakly influence variables otherwise treated as exogenous. We study reinforcement learning under bounded violations of exogeneity, where each action perturbs the exogenous transition law. The perturbation is controlled by an impact parameter . We introduce two optimistic algorithms for this perturbation: a tabular method and Exogenous Value-Targeted Regression (Exo-VTR), which exploits linear-mixture structure in large exogenous spaces. We prove high-probability regret of the form , where is the number of episodes, retaining the statistical benefit of exogenous structure while paying an explicit perturbation penalty. In the tabular setting, no-bonus predict-then-optimize (PTO) attains the same regret order. We also provide lower bounds on first-order loss for specified exogenous planner classes. A separate two-stage construction establishes first-order regret for Exo-UCBVI-, PTO, and Exo-VTR- with a one-dimensional reference representation. Experiments validate the execution-family separation and probe the roles of clipping and optimism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.