acceptodds
Under review as a conference paper at ICLR 2027

Power Mirror: Offline Policy Improvement with Bounded Reweighting

Abstract

Offline reinforcement learning (RL) must balance policy improvement with reliable use of logged data. We introduce *Power Mirror*, a policy-improvement framework derived from shifted power-mean value backups. With a suitable shift, the induced updates satisfy explicit two-sided bounds on actionwise probability ratios, together with population improvement guarantees and sensitivity bounds for critic error. We develop two complementary algorithms: fixed-anchor value iteration maintains a permanent bound relative to the behavior policy, while moving-anchor policy iteration improves monotonically and converges in value to the support-constrained optimum under exact evaluation and persistent updates. We also construct a ratio-preserving actor for continuous actions and a held-out certification rule that deploys a learned policy only when improvement is certified, otherwise retaining the behavior policy. Together, these results connect value estimation, policy optimization, and deployment through explicit control of policy reweighting. Experiments support our theoretical findings, with Power Mirror achieving strong policy improvement while remaining conservative under limited data coverage

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.