acceptodds
Under review as a conference paper at ICLR 2027

Where to Query: Dataset Aggregation as Quadrature in Interactive Imitation Learning

Abstract

Interactive imitation learning often collects data under mixtures of expert and learner behaviour to mitigate covariate shift, but mixture points and weights are typically specified as a training schedule without an explicit approximation target. We show that, along a fixed expert–learner path, weighted dataset aggregation is a numerical quadrature rule for the path-average occupancy. This target assigns positive mass to every state reachable under some mixture on the discounted path, whereas aggregates supported only on the endpoint policies can miss states that occur under interior mixtures. For finite-state Markov Decision Processes (MDPs), classical quadrature theory yields schedule-dependent rates, including an endpoint-Riemann bound and geometric convergence for Gauss–Legendre, with constants that grow with the effective horizon or mixing time. Exact evaluation reproduces the rate hierarchy and, on Garnets, the predicted endpoint limit. Under independent finite sampling, Gauss–Legendre and midpoint reach an error no more than 20% above that of direct path-average sampling with the same budget in at most three rounds at every tested budget for the median instance, whereas endpoint Riemann, DAgger, and random-uniform schedules require many more rounds or fail to reach this level as the budget grows. In continuous control, with retraining held fixed, path-covering node sets lower reference-path imitation loss relative to DAgger-P's endpoint nodes on three of four environments but improve return significantly on at most two.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.