UserScaler: Beyond Environment Scaling—Unlocking the User Dimension for Agent Training
Abstract
Scaling the training of large language model (LLM) agents relies on diverse and reliable environments. Existing methods leverage automatic pipelines to synthe- size interactive, stateful, and scalable environments. Nevertheless, they overlook the substantial gap between idealized users and real-world users: realistic users often hold partial intents, exhibit variable phrasing styles, and react dynamically as interactions proceed. This discrepancy hinders the effectiveness of existing agents in real-world application. To tackle this issue, we propose UserScaler, which jointly generates plausible users and environments to bridge the gap be- tween agent training and real-world deployment. UserScaler consists of three core components: (1) Realistic user-centric environment synthesis: a search agent constructs each environment following specifications of real services and derives user personas from market research of the target service; (2) Realistic user-driven task generation: user requests are synthesized conditioned on user personas and sampled toolchains; (3) User-guided trajectory collection: during rollout, simu- lated users actively engage in the interaction, whose behaviors are guided by an automatically derived rubric covering both process and outcome. We instanti- ate 209 environments (109 single-service and 100 cross-domain compositions), each equipped with roughly 50 user personas, yielding 12,073 training trajecto- ries. Training Qwen3 models from 1.7B to 8B on this data consistently improves pass rates over corresponding baselines trained with EnvFactory, AWM, and En- vScaler across three agentic benchmarks: the 8B model reaches 44.3% on τ2 -bench, 19.8% on VitaBench, and 11.3 on UserBench. Under a matched budget of 6,474 trajectories, user-guided rollouts outperform direct ones, confirming that structured user modeling provides signal beyond additional data volume. We re- lease the environments, training data, and models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.