Dynamic Elite Replay: Learning Beyond Suboptimal Experts in Human-in-the-Loop Reinforcement Learning
Abstract
Human-in-the-loop reinforcement learning (HIL-RL) uses expert intervention and demonstrations to improve training safety and learning efficiency. Many existing methods use expert behavior as an important source of positive supervision, which is effective when reliable expert guidance is available. However, when the expert is suboptimal, continued reliance on such supervision can bias policy learning toward the expert's behavior and make it difficult for the agent to fully exploit better behaviors discovered through its own exploration. To address this problem, we propose Dynamic Elite Replay (DER), an intervention learning framework that integrates elite sample selection and weighted elite learning, enabling the agent to fully exploit high-quality behaviors discovered through its own exploration. DER first identifies high-quality agent-generated experience according to task performance and safety. When complete successful trajectories are scarce, it further extracts safe progress segments that capture useful local improvements. During replay, retained elite samples are re-evaluated using current Q-value advantage, twin-critic agreement, sample freshness, and disagreement with the supervisor, reducing the influence of outdated or unreliable experience. The reweighted elite experience is then used for behavior-cloning and critic updates to reinforce validated high-quality agent behaviors. Experiments in MetaDrive show that DER can achieve performance above a fixed suboptimal supervisor while reducing safety failures. These results suggest that expert intervention can guide training without necessarily defining the final performance ceiling when high-quality behaviors discovered by the agent are effectively identified and reused.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.