Enhancing Reinforcement Learning for Autonomous Driving with Outcome-Related Reward Shaping
Abstract
Designing effective reward functions for autonomous driving remains a major challenge since high-level human driving concepts are difficult to translate into precise numerical signals. While current approaches attempt to address this using expert-designed rewards or semantic signals from language models, these approximations often fail to align with the actual physical outcomes of the driving task. To address this, we propose a reward shaping method utilizing empirical environment outcomes as the evaluation target. Our approach introduces Outcome-Derived State Evaluation (ODSE), a state evaluator learned from reference trajectories by using trajectory-level terminal outcomes as supervised targets for visited states. By integrating these evaluations through the potential-difference formulation of Potential-Based Reward Shaping (PBRS), we derive a dense shaping signal that improves sample efficiency while only adding a controlled, explicitly characterized outcome term to the learning objective. Evaluations on the VLM-RL benchmark show that our method outperforms state-of-the-art reward models, thereby demonstrating that outcome-related signals can successfully enhance training effectiveness and establish a more robust driving strategy across diverse scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.