acceptodds
Under review as a conference paper at ICLR 2027

A Unified Framework for Learning, Stabilization, and Optimization in Online Nonlinear Control

Abstract

We study online stochastic optimal control with unknown nonlinear dynamics. We address three limitations of existing approaches: prespecified parametric dynamics; reliance on a priori bounded trajectories or a known stabilizing controller; and restrictive noise assumptions, such as bounded disturbances or known covariance structures. Specifically, we propose Gibbs-Randomized Model-Based Adaptive Control (), a two-stage algorithm that first learns and certifies a stabilizing feedback gain and then improves control performance through sieve-based dynamics estimation, noise-covariance estimation, and residual-policy updates. Under mild assumptions, we derive finite-sample bounds for the dynamics and covariance estimation errors and establish regret guarantees relative to the optimal finite-horizon controller. We further specialize the theoretical analysis to linear quadratic regulator, nonlinear Fourier dynamics, and nonlinear transient price impact. Under the corresponding conditions, the proposed controller achieves a sublinear regret bound that scales as the square root of the time horizon and is asymptotically optimal.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.