ReAgent: Rethinking Agent Training through Scenario Co-Design
Abstract
Agentic reinforcement learning scales training by expanding programmatic environments and synthesizing diverse tasks. Yet this scaling can paradoxically narrow training coverage through how data are constructed. To make such synthesis tractable at scale, existing pipelines often derive tasks from executable paths or successful workflows, thereby inducing a construction-induced selection bias toward feasible requests. Real users, however, state goals without full visibility into the environment state or its constraints, so agents must identify conflicts, guide users toward feasible revisions, and refuse when necessary. Such resolution behaviors cannot be captured by task synthesis alone, as conflicts and valid resolutions are co-constructed by environment constraints, user intent and appropriate reward. This interdependence motivates us to rethink agent training by shifting the unit of construction from feasible tasks alone to \itshape the whole training scenario (i.e., task, user, and reward together). We call this principle scenario co-design and instantiate it through ReAgent. Starting from conflicts and resolutions grounded in the environment, ReAgent jointly constructs the task request, simulated user behavior, and reward criteria as a coherent conflict scenario, enabling scalable training with broader coverage of feasible, blocked, and partially resolvable requests. Empirically, augmenting training with these co-designed conflict scenarios not only directly improves Qwen3-8B and Qwen3-30B-A3B-Thinking-2507 on constraint-governed agent benchmarks such as -Bench but also even enables them to outperform competitive baselines across general agent benchmarks. These results suggest that broadening training coverage within environments through scenario co-design offers a promising direction for scaling agent training beyond simply expanding environments and feasible tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.