UnsafeArena: Agent and Environment Co-Evolution for Safety and Utility
Abstract
Training agents in synthetic environments can improve complex task performance, yet scalable interactive environments for agentic safety reinforcement learning remain scarce. We introduce UnsafeArena, a framework that jointly improves agent utility and safety by co-evolving policies and executable environments grounded in real-world settings. UnsafeArena combines challenging utility tasks with safety tasks involving contextual risks, supported by reliable verifiers for reinforcement learning. After each training round, dynamic evaluation across difficulty levels guides targeted task generation and incremental environment refinement. We further introduce Compass-GRPO, a trajectory-calibrated extension of group relative policy optimization that incorporates dense self-distillation guidance into policy advantages. A self-teacher compares token probabilities with and without task hints along the unhinted student's own trajectories. This guidance is jointly calibrated against utility and safety to support learning when trajectory rewards provide little differentiation. We instantiate UnsafeArena with 2,661 stateful environments, construct over 40,000 tasks, and conduct large-scale reinforcement learning. Experiments demonstrate substantial improvements over the Qwen3-14B base model: UnsafeArena-14B reduces AgentSecurityBench attack success by 59.5 percentage points and raises -Bench task success by 8.1 points. Gains generalize to external benchmarks without evaluation hints. These results show that environment synthesis and targeted policy optimization can develop agents that accomplish complex tasks while respecting execution boundaries.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.