AOS-Bench: Can LLM Agents Manage Dynamic Workloads?
Abstract
Coding agents are increasingly asked to handle multiple requests on the same computer, yet current benchmarks evaluate tasks one at a time. We introduce AOS-Bench, a benchmark that evaluates a persistent agent system as a workload manager: tasks arrive over time and must be delivered under shared CPU and memory limits. Across five model–harness pairs, isolated task success poorly predicts workload delivery. On medium workloads, systems differ by 13 percentage points in isolated success but 29 points in stream success; on heavy workloads, two systems with identical isolated baselines deliver 72.2% and 49.7% of the same tasks. Execution-policy interventions trace much of this gap to task intake: serial dispatch leaves 73% of heavy-workload tasks unclaimed. Yet more elaborate management does not necessarily help—a fixed concurrency of two matches explicit management guidance on medium workloads, while no tested policy uses more than 30% of the provisioned CPU capacity. These results identify workload management as a distinct challenge for agents and establish AOS-Bench as a way to measure it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.