acceptodds
Under review as a conference paper at ICLR 2027

AReward: Synthesizing Real-World Environments with Executable State Rubrics

Abstract

Language agents learn through interaction with stateful environments, but sparse outcome rewards provide limited feedback on intermediate progress. Rubric-based rewards offer finer-grained evaluation, yet their criteria must reflect user requirements and be checked reliably against execution. We introduce AReward, which jointly synthesizes transactional environments, tasks, and executable state rubrics through a shared formulation of create, read, update, and delete (CRUD). Tool execution determines state changes, and task requirements specify the conditions that count as completion. Reference solution replay verifies task reachability, while reward auditing examines the consistency of acceptance criteria with user requirements. Evaluating these rubrics after each action exposes progress and reversals in rubric satisfaction. STG-GRPO, our state-transition graph variant of Group Relative Policy Optimization, uses this evidence and progress along the reference state chain to refine trajectory comparisons. Combined with supervised initialization, reinforcement learning with a staged curriculum, and long-tail task management, AReward provides a training framework in which reward criteria are tied to user requirements and their evaluation is grounded in actual execution. Code is available at https://anonymous.4open.science/r/stg-anonymous-248B/

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.