acceptodds
Under review as a conference paper at ICLR 2027

Plug-and-Play Graph-Based Intrinsic Reward for Exploration of Cost-Aware Reinforcement Learning

Abstract

Introducing action costs increases the difficulty of exploration in complex, sparse-reward state spaces for reinforcement learning (RL). To address this challenge, we propose a novel graph-based reward shaping method, termed Graph Intrinsic Reward (GIR), which uses graph to model the structure of explored state space to enable efficient exploration-exploitation in cost-aware RL. Specifically, GIR incrementally constructs a graph in which each node maintains both visit counts and value estimates, from which a count-based exploration reward and a potential-based exploitation reward are derived to guide the agent in escaping initial suboptimal states under costly actions. Through reward-gated modulation, GIR reduces ineffective exploration, while potential-based value propagation over the graph guides the agent toward high-return regions. We further analyze the impact of GIR on policy optimality and show that, from a theoretical perspective, GIR asymptotically preserves the optimal policy. Extensive experiments demonstrate that the proposed reward shaping strategy effectively guides exploration towards goal states and outperforms state-of-the-art exploration methods in cost-aware environments. Our implementation is available at https://anonymous.4open.science/r/GIR-48BD/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.