Fast Reward Evaluation in Reinforcement Learning via Simulation Optimization
Abstract
Reinforcement learning (RL) plays an increasingly critical role in modern artificial intelligence, including decision-making tasks and large language model (LLM) training. In RL, the reward function determines what behavior the policy is optimized to learn. In practice, different candidate rewards can lead to substantially different policy performance, making reward evaluation crucial for selecting rewards that effectively guide policy optimization. However, evaluating candidate rewards often requires training a separate policy for each candidate, resulting in substantial computational cost and making reward evaluation a major bottleneck. To address this challenge, we reframe reward evaluation as a simulation optimization problem in operations research, which studies efficient selection under limited budgets. Under this formulation, reward evaluation is cast as a sequential budget allocation problem. We then propose Fast Reward Evaluation by adapting Optimal Computing Budget Allocation (OCBA), a representative simulation optimization method, to guide the allocation of training budget across candidate reward functions. Empirical results on sparse-reward RL, automatic reward generation, and LLM reasoning tasks demonstrate that evaluation efficiency is significantly improved, and that high-quality reward functions are more reliably identified under the same computational budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.