SARC: Interpretable State-Adaptive Reward Composition for Sparse Reward Reinforcement Learning
Abstract
Sparse-reward reinforcement learning suffers from severe credit assignment challenges, as agents often receive no informative feedback within an episode. Reward decomposition and shaping methods alleviate this problem but often introduce redundant information or reward hacking. Recent latent-reward approaches provide structured, multi-dimensional signals, but their composition does not necessarily expose how individual factors influence policy behavior across states. We propose State-Adaptive Reward Composition (SARC), which composes the action reward as a linear combination of the environment reward and multi-dimensional latent rewards. The composition weights are produced by a state-dependent intention function trained through bilevel optimization: the upper level optimizes the intention function to maximize environment return, while the lower level optimizes the policy to maximize the composed return. The learned intention weights explicitly and interpretably reveal the agent's task priorities at each state. We further introduce stabilization techniques for this bilevel training. Across eight sparse-reward continuous-control tasks, SARC achieves strong performance against competitive baselines. Intervention experiments show that the learned weights causally steer behavior, while ablations characterize the contributions of the main components.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.