R-Split: Population-Level Reasoning over Reward Mechanisms for Automated Reward Design with LLMs
Abstract
Reward functions strongly affect how reinforcement learning (RL) agents learn, but designing them for complex control tasks requires substantial expert effort. Large language models (LLMs) can generate and improve reward code. However, existing methods typically apply LLM reasoning to individual candidates, which limits their use of population-wide information for candidate comparison and search planning. We propose R-Split, which analyzes candidate reward designs, compares them across the population, makes evolution decisions, and generates new reward code. R-Split summarizes each evaluated reward function as a structured set of reward mechanisms: textual descriptions of how its components shape learning, grounded in reward code and training-process data. A similarity filter removes candidates with highly similar mechanism descriptions, reducing repeated exploration. We evaluated our approach against multiple baselines on ten low-level control tasks, using a shared PPO evaluation pipeline and five independent runs for each selected reward. R-Split achieves competitive performance across both locomotion and manipulation tasks. These results show that comparing reward mechanisms across a population can help improve LLM-based reward design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.