acceptodds
Under review as a conference paper at ICLR 2027

LOOPADAPT: PERSISTENT POLICY ADAPTATION VIA LOOP-TO-SCOPE CERTIFICATION

Abstract

Policies adapted over long deployment horizons face a challenge beyond recovering from the current failure: every update also changes the adaptation options available in the future. We formalize this setting as persistent policy adaptation and introduce LOOPADAPT, a framework that links recurrent failure modes to the parameter scopes in which they can be jointly corrected. LOOPADAPT models low-progress behavior as recurrent loops, assigns each retained loop a differentiable objective, and tests whether their gradients admit a common descent direction within a prescribed update scope. This yields a loop-to-scope certificate that identifies first-order incompatibilities among recurrent failures and, when feasible, guarantees finite-step improvement while accounting for gradient estimation error, local curvature, and update radius. LOOPADAPT then ranks certified candidates by long-horizon value and validates them before commitment, explicitly accounting for how present updates reshape future checkpoints and adapter choices. Across 128 transient kernels, the certified direction improves every retained loop in all eight-dimensional settings, compared with 62.5% for mean descent and 96.875% for PCGrad. In PointMass, incorporating the loop objective reduces mean loop potential by 41.6% while preserving 100% task success. Across ten 300-window streams, long-horizon and immediate update choices differ on 16.1% of windows, showing that locally attractive updates can alter future adaptation opportunities. These results show that persistent policy adaptation is fundamentally a joint problem of recurrent-failure geometry, scope-constrained update feasibility, and cross-window value.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.