acceptodds
Under review as a conference paper at ICLR 2027

Hybrid Reinforcement Learning under Dynamics Shift: Policy Optimization, Certification, and Adaptation

Abstract

Hybrid reinforcement learning combines offline data with online interaction. When these data come from environments with different transition dynamics, the learner must decide how to use the source and when to rely on target observations. We study policy optimization, target certification, and adaptation under such dynamics shift. For general function classes satisfying realizability and Bellman completeness in both environments, we analyze pessimistic policy optimization with source-consistent critics. Its guarantees depend on target Bellman residuals along the learned policies and a target-optimal comparator, quantified through policy-dependent transfer moduli. A block-based algorithm uses its own target observations to constrain the deployed-policy residual; a separate comparator term and the cost of failed tests remain. An incompatibility test supports a handoff to online learning, while a return-based master adapts between source-informed copies and an online baseline, with coefficient-dependent overheads. Structured upper bounds separate source uncertainty from dynamics shift in tabular, linear, and block MDPs. Lower bounds show that the resulting square-root modulus dependence is necessary on hard families when target observations cannot distinguish the competing models. In finite-class experiments, an optimal source policy can trigger handoff while a better unvisited route remains undetected, and return-based selection partly recovers from a misleading source at a substantial cost relative to target-only controls.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.