acceptodds
Under review as a conference paper at ICLR 2027

Solving Sparse-Reward Environments using Meta-Gradients

Abstract

Exploration is essential in reinforcement learning since agents rely on trial and error to learn an optimal policy. However, when rewards are sparse, naive exploration strategies, like noise injection, are often insufficient. Intrinsic rewards can provide principled guidance for exploration by combining them with extrinsic rewards to optimize a policy or using them to train subpolicies for hierarchical learning. However, the former approach induces suboptimality as it changes the performance objective, while the latter is sample inefficient and also suboptimal. To address these challenges, we propose an algorithm, intrinsic reward policy optimization (IRPO), that uses intrinsic rewards within a meta-gradient framework to drive exploration without changing the performance objective. We achieve this by updating a shared meta-policy such that one of the subpolicies (derived from the meta-policy using intrinsic rewards) optimizes extrinsic rewards. We demonstrate that IRPO outperforms relevant baselines in sparse-reward environments with similar wall-clock time to standard policy gradient methods, and justify IRPO design choices through ablation studies. We also prove that IRPO converges to an -approximate first-order stationary point.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.