Power Mirror: Offline Policy Improvement with Bounded Reweighting
Abstract
Offline reinforcement learning (RL) must balance policy improvement with reliable use of logged data. We introduce *Power Mirror*, a policy-improvement framework derived from shifted power-mean value backups. With a suitable shift, the induced updates satisfy explicit two-sided bounds on actionwise probability ratios, together with population improvement guarantees and sensitivity bounds for critic error. We develop two complementary algorithms: fixed-anchor value iteration maintains a permanent bound relative to the behavior policy, while moving-anchor policy iteration improves monotonically and converges in value to the support-constrained optimum under exact evaluation and persistent updates. We also construct a ratio-preserving actor for continuous actions and a held-out certification rule that deploys a learned policy only when improvement is certified, otherwise retaining the behavior policy. Together, these results connect value estimation, policy optimization, and deployment through explicit control of policy reweighting. Experiments support our theoretical findings, with Power Mirror achieving strong policy improvement while remaining conservative under limited data coverage
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.