Evolution Only When Needed: Triggered Evolutionary Data Injection for Off-Policy Reinforcement Learning
Abstract
Evolutionary reinforcement learning (ERL) can expand exploration for off-policy learners, but maintaining a population throughout training requires repeated auxiliary rollouts even when the learner is still improving. Moreover, indiscriminately inserting population-generated trajectories into replay can shift the replay distribution away from the current learner and degrade off-policy updates. We propose Triggered Evolutionary Data Injection (TEDI), which treats evolutionary assistance as an on-demand replay-data intervention rather than a persistent policy-optimization process. TEDI jointly decides when to invoke evolutionary exploration, which generated experience should enter replay, and how much influence each intervention may have. When persistent learning stagnation is detected, TEDI creates a temporary population of parameter-perturbed policy probes, collects their trajectories outside the main replay buffer, filters candidate experience at both transition and trajectory levels, and injects only a bounded, balanced subset before discarding the population. Thus, evolutionary search affects learning only through selectively admitted replay data. Across four MuJoCo continuous-control tasks, TEDI achieves the highest mean peak return in comparisons with six baselines. Ablation studies show that triggering, filtering, and bounded injection each contribute to performance, while runtime comparisons show that TEDI has the lowest average training time among the compared ERL methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.