Where to Query: Dataset Aggregation as Quadrature in Interactive Imitation Learning
Abstract
Interactive imitation learning often collects data under mixtures of expert and learner behaviour to mitigate covariate shift, but mixture points and weights are typically specified as a training schedule without an explicit approximation target. We show that, along a fixed expert–learner path, weighted dataset aggregation is a numerical quadrature rule for the path-average occupancy. This target assigns positive mass to every state reachable under some mixture on the discounted path, whereas aggregates supported only on the endpoint policies can miss states that occur under interior mixtures. For finite-state Markov Decision Processes (MDPs), classical quadrature theory yields schedule-dependent rates, including an endpoint-Riemann bound and geometric convergence for Gauss–Legendre, with constants that grow with the effective horizon or mixing time. Exact evaluation reproduces the rate hierarchy and, on Garnets, the predicted endpoint limit. Under independent finite sampling, Gauss–Legendre and midpoint reach an error no more than 20% above that of direct path-average sampling with the same budget in at most three rounds at every tested budget for the median instance, whereas endpoint Riemann, DAgger, and random-uniform schedules require many more rounds or fail to reach this level as the budget grows. In continuous control, with retraining held fixed, path-covering node sets lower reference-path imitation loss relative to DAgger-P's endpoint nodes on three of four environments but improve return significantly on at most two.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.