Uncertain Rewards Bring More Possibilities: Bidirectional Reward Noise for Exploration in Reinforcement Learning
Abstract
Deep reinforcement learning (DRL) relies on reward-driven policy updates, yet sufficient exploration remains a critical challenge for discovering high-quality policies. Due to the tendency to pursue high rewards, agents are prone to becoming trapped around rewarding states encountered early in training, hindering exploration of potentially more valuable regions and leading to premature convergence. Existing exploration methods commonly introduce additional stochasticity or intrinsic incentives to encourage broader exploration. However, stochasticity alone lacks explicit directional control, while intrinsic rewards often require additional modeling components. In this paper, we propose Bidirectional Reward Noise (BiRN), a lightweight exploration mechanism that directly perturbs training rewards. By injecting bidirectional noise, BiRN achieves two complementary effects: negative noise mitigates excessive attraction toward dominant rewarding transitions, while positive noise strengthens incentives for transitions with positive learning potential. Extensive experiments show that BiRN consistently improves final performance across diverse RL benchmarks with minimal additional overhead, with particularly strong gains in challenging sparse-reward settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.