acceptodds
Under review as a conference paper at ICLR 2027

AOS-Bench: Can LLM Agents Manage Dynamic Workloads?

Abstract

Coding agents are increasingly asked to handle multiple requests on the same computer, yet current benchmarks evaluate tasks one at a time. We introduce AOS-Bench, a benchmark that evaluates a persistent agent system as a workload manager: tasks arrive over time and must be delivered under shared CPU and memory limits. Across five model–harness pairs, isolated task success poorly predicts workload delivery. On medium workloads, systems differ by 13 percentage points in isolated success but 29 points in stream success; on heavy workloads, two systems with identical isolated baselines deliver 72.2% and 49.7% of the same tasks. Execution-policy interventions trace much of this gap to task intake: serial dispatch leaves 73% of heavy-workload tasks unclaimed. Yet more elaborate management does not necessarily help—a fixed concurrency of two matches explicit management guidance on medium workloads, while no tested policy uses more than 30% of the provisioned CPU capacity. These results identify workload management as a distinct challenge for agents and establish AOS-Bench as a way to measure it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.