acceptodds
Under review as a conference paper at ICLR 2027

The Improvement Guarantee of Learning MPC Leaks on Nonlinear Systems, and How to Bound It

Abstract

Learning Model Predictive Control (LMPC) improves at a repetitive task by storing executed trajectories and reusing their realized cost-to-go as terminal cost. Its improvement guarantee is proved on the stored states, but a nonlinear closed loop only ever reaches a tolerance-neighbourhood of them, so the guarantee must survive whatever extends the value off the store. We define the consistency leak, the amount by which one step of true progress can fail to buy one step of promised value, and prove that the guarantee degrades by at most the leak per step. Measured on stores that a minimum-time LMPC loop generates for itself, the position-indexed constructions the literature runs (convex safe sets, nearest-neighbour lookups, a learned value) leak more as experience accumulates: within nine iterations the improvement bound exceeds the cost it bounds. State-aware constructions keep the leak bounded but expose almost none of the stored improvement. The remedy is to take the nearest neighbour within each stored trajectory and then the minimum over trajectories; its leak is bounded by plant and design constants alone. In closed loop the per-trajectory minimum completes all 176 iterations across two plants while the pooled lookup completes fewer than half, at 0.58-2.2x the cost per MPC solve. Code is available at https://anonymous.4open.science/r/lmpcleak-54F8.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.