DESIGN THEFEEDBACK, NOTJUST THEREWARD: ENVIRONMENTS FOROPEN-ENDED REINFORCEMENT LEARNING
Abstract
Reinforcement learning has advanced large language models on tasks with reliableverifiers, where checking an outcome provides a clear basis for reward. Extendingthis approach to open-ended tasks requires environments that can both evaluatecomplex, partially successful outputs and turn these judgments into effective learn-ing guidance. We propose a reinforcement learning framework built around atraining environment that integrates human domain knowledge with model reason-ing. Human domain knowledge, encoded in rubrics, anchors evaluation to taskintent, while model reasoning combined with hierarchical rewards distinguishescomplex open-ended outputs and links diagnosed failures to likely responsible out-put regions. Furthermore, these diagnostic attributions guide token-level weights inthe GRPO objective, thereby translating the ambiguous notion of success in open-ended tasks into concrete guidance for policy optimization. We instantiate thisframework in interactive web application generation, a representative open-endedgeneration task. The resulting model achieves an average pass rate of 41.25% onMiniAppBench and 76.19% on ArtifactsBench. Ablation studies show the contri-butions of individual components, while additional out-of-distribution evaluationssupport transferable gains without degrading performance on other evaluated tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.