CloudGym: Can AI Agents Manage Shared Cloud Infrastructure?
Abstract
Managing cloud infrastructure requires substantial expertise and operational effort, motivating automation with AI agents. As cloud infrastructure is inherently shared, cloud management agents must coordinate with other participants whose actions can invalidate their plans or block their operations. Existing evaluations largely focus on static plans or isolated execution, leaving coordination in shared environments insufficiently tested. We present CloudGym, a benchmark for evaluating cloud management agents in live, shared cloud infrastructure, with 102 cases spanning 19 AWS services and 65 resource types. Each case introduces controlled concurrent participants whose actions create incompatible requirements, ambiguous targets, or temporary execution conflicts. Semantic oracles use initial and final cloud states and the record of activated participants to check whether outcomes satisfy the task intent under a specified resolution policy. To support reliable evaluation, an automated pipeline generates these cases from existing management tasks and certifies their behavior through execution against AWS. Our evaluation shows that frontier agents perform poorly on CloudGym, with the top-performing model, Opus 5.5, achieving an overall pass rate of 34.3% compared with 97.1% in isolated execution. This gap highlights coordination as a major obstacle to agentic cloud management. We open-source the CloudGym dataset and evaluation framework to enable future research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.