Data-Driven Hamiltonian for Neural Optimal Control
Abstract
Reinforcement learning (RL) has emerged as a prominent data-driven approach to optimal control, but often requires substantial amounts of data and provides limited guarantees when encountering out-of-distribution state-action pairs. In this paper, we propose an alternative approach grounded in the well-established Hamilton-Jacobi framework in optimal control theory. Rather than requiring an explicit analytical model of the system dynamics to evaluate the Hamiltonian, we replace this with a data-driven Hamiltonian (DDH), which we generalize to a broad class of optimal control problems. In contrast to standard RL approaches, the DDH provides a principled mechanism, requiring only minimal prior system information such as a Lipschitz constant, for quantifying and accounting for uncertainty in regions not covered by the data, enabling the derivation of controllers with performance guarantees. Additional prior system knowledge, when available, can be readily incorporated to improve performance. For high-dimensional systems, we further introduce a physics-informed neural approximation of the value function, trained using a novel procedure that generates worst-case fictitious rollouts evaluated through the DDH. On an LQR problem, the DDH reaches near-optimal cost on every dataset we test, including random data on which offline RL methods remain far from optimal, and needs far fewer samples than online RL. On a nonlinear lunar lander benchmark, the DDH achieves an optimality gap to a trajectory-optimization reference more than smaller than CQL and FQL, and matches the performance of SAC with approximately fewer samples. These results establish the DDH as a promising framework for data-driven optimal control, particularly in settings where available data are limited in quantity or quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.