acceptodds
Under review as a conference paper at ICLR 2027

Agentic DAgger: Self-Improving Interactive Imitation Learning via Agentic Harness

Abstract

Interactive imitation learning (IIL), such as DAgger, improves a policy by collecting expert feedback on learner-visited states. However, conventional IIL typically relies on a static feedback-generation pipeline comprising a likely suboptimal expert and training experience with diminishing learning value as the policy improves. Consequently, further policy improvement can be bottlenecked by both the quality of expert feedback and the utility of accumulated experience. We introduce Agentic DAgger, a training harness that jointly improves feedback quality and the learning value of experience alongside the learner. Agentic DAgger employs two specialized agent roles operating through verifier-grounded agentic loops. Within each training round, a feedback-evolution loop diagnoses original expert proposal, propose refined corrective labels, and verifies the labeling quality. Between rounds, an experience-evolution loop analyzes the learner-expert gap on the current training experience and acquires novel training experiences that maximize learning value. Together, these loops transform conventional IIL into a self-improving training process that progressively improves corrective feedback and the learning value of training experience. We evaluate Agentic DAgger by improving pretrained policies on NAVSAFE-∞, a verifiable closed-loop driving benchmark. Our experiments show that Agentic DAgger achieves higher closed-loop driving success than conventional IIL baselines under matched training budgets. It further improves learner performance across successive IIL rounds, demonstrating sustained policy self-improvement. These findings demonstrate the effectiveness of Agentic DAgger and highlight its potential to sustain policy self-improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.