Cost-Aware Active Learning of Optimal Control Policies on the Switching Set
Abstract
We consider discounted optimal control of control-affine systems under input constraints. The optimal feedback of such problems saturates and switches between input levels across surfaces in the state space. Imitating such a controller requires one optimal control solve per expert sample, so the samples must be placed where they matter. For the policy over a finite input grid, learned as a classifier, we show that these samples should lie on the switching surfaces. Their value there is set by the cost slope g, the rate at which the cost of the wrong input grows with the distance from the surface. We prove that the cost-sensitive loss of the learned policy is the -weighted squared error of its switching surfaces and that the optimal sample density on them is proportional to , with the local rate of the learner. Margin sampling loses a constant factor, computable from the expert before learning, and an idealised learner improves the loss rate of uniform sampling from to in state dimensions. The practical sampler estimates from the already-returned action values, with no additional expert samples. On the inverted pendulum, two mountain-car configurations and a four-dimensional cart-pole, the sampler matches the accuracy of uniform sampling with to more than times fewer expert samples and its closed-loop cost gap with to times fewer. The factor computed beforehand predicts on which systems the cost weighting pays off.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.