acceptodds
Under review as a conference paper at ICLR 2027

Teacher-Policy Value Identification from Control Demonstrations with Teacher-Relative Performance Guarantees

Abstract

SDRE controllers are useful teachers for high-dimensional nonlinear systems, but their algebraic value surrogate need not coincide with the infinite-horizon cost-to-go realized by the resulting policy. We propose a control-supervised approach to learn this teacher-policy value from held-input transitions and control labels, without value labels or statewise drift evaluations during training. A single neural value function, trained by one-step policy-value consistency and control supervision, yields both an estimate of the teacher-policy value and a gradient-defined controller that avoids per-state Riccati solves at deployment. We prove that a uniform one-step residual identifies the realized cost-to-go; Through a completion-of-the-square identity, value consistency and control supervision bound the HJB defect. For bounded finite-energy trajectories, the deployed cost and control deviation are bounded by the teacher cost plus residual/discretization errors, without optimal references or pointwise imitation. LQR anchoring ensures local exponential stability for all network parameters, independent of training accuracy. The method eliminates drift knowledge, accelerates SDRE deployment, and offers provable value recovery, HJB consistency, and stability. Experiments on 40D Cucker–Smale, 51D Allen–Cahn, and 64D Burgers confirm stable closed-loop performance and rollout-value recovery while replacing online Riccati solves with amortized neural feedback.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.