acceptodds
Under review as a conference paper at ICLR 2027

Arbitrated Handover from Baseline Control to Reinforcement Learner

Abstract

A running system cannot pause for a new controller to learn. Online controller adaptation avoids the pause by keeping a fixed baseline in control while a reinforcement learning policy, the learner, trains on the live system. Gating blends the two actions, passing control gradually from baseline to learner over training. The gate must judge when the learner is skilled enough to take over; handing over too early costs performance and can cause failures the baseline would have avoided. We identify two core challenges in this setting. The first is that improving the learner and judging the handover need different value estimates: one of the learner acting alone, one of the blend that actually runs. The second is that, because the learner does not act on its own until the end of training, its value estimate carries extrapolation error. We find that neglecting either results in premature handover, causing performance degradation and task failure during training. We propose arbitrator-gated handover: an arbitrator judges the executed blend with value estimates of its own, and we apply techniques to limit the effects of value estimation error in this setting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.