Wait or Take Over: Forecasting OOD Persistence and Recovery for Efficient Human-in-the-Loop Imitation Learning
Abstract
Robot-gated interactive imitation learning allows an agent to request human intervention and learn from the resulting corrective demonstrations. Deciding when to request human takeover is central to using limited human effort effectively. Existing intervention mechanisms commonly rely on pointwise novelty, uncertainty, or disagreement, leaving an ambiguity: a novice outside human support may recover autonomously or continue along a persistent deviation. Intervening in the former case wastes limited human effort. We formulate robot-gated interactive imitation learning as a sequential intervention decision problem and develop Recovery-aware Gating for Interactive Imitation Learning (ReGIL). Drawing on Hamilton–Jacobi reachability analysis, we learn a support-persistence value that estimates whether the novice can remain within human support under its own policy. A short-horizon dynamics model predicts a return to support, while the persistence value assesses behavior after that return. A sequential recovery assessment combines these predictions with observed changes in persistence to determine when human intervention is warranted. This allows the novice to continue when autonomous recovery remains plausible and requests human takeover when predicted recovery is insufficient and observed persistence declines. Experiments in MetaDrive and MetaWorld demonstrate the effectiveness of the framework in learning autonomous policies from limited human intervention. The corresponding source code is available at [https://anonymous.4open.science/r/ReGIL-7864](https://anonymous.4open.science/r/ReGIL-7864).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.