HAIL: Heterogeneity-Aware Learning from Human Interventions with Suboptimal Experts
Abstract
Human-in-the-loop reinforcement learning (HITL-RL) has recently attracted increasing attention as a means to improve the learning efficiency and safety of reinforcement learning in complex tasks by incorporating human expert interventions. However, existing methods often implicitly assume that human experts can consistently provide high-quality guidance during training. In practice, human experts are often suboptimal, and the supervision they provide exhibits both intervention-level heterogeneity and temporal-level heterogeneity. To address these issues, we propose HAIL (Heterogeneity-Aware Learning from Human Interventions with Suboptimal Experts), a framework that explicitly models the heterogeneity of human supervision provided by suboptimal experts. Specifically, Quality-Weighted Behavioral Learning (QWBL) estimates the relative quality of different intervention behaviors and dynamically adjusts their optimization weights, allowing intervention samples with different quality levels to contribute differently to policy optimization. Effective-Horizon Retention (EH) selects preference information with higher learning value from different predicted future states while retaining part of the remaining candidates through sparse tail sampling, reducing the interference from preference information with lower learning value while maintaining trajectory-level supervision coverage. Experimental results in autonomous driving and continuous control environments demonstrate that HAIL improves policy performance with suboptimal human experts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.