acceptodds
Under review as a conference paper at ICLR 2027

Schedule-Overcooked: Multi-Task Allocation for LLM Agents under Multiple Constraints

Abstract

Large language model (LLM) based multi-agent systems increasingly execute concurrent tasks over a shared pool of agents and tools, making the allocation of dependent subtasks a central determinant of system efficiency. Yet allocation is evaluated only through end-to-end task success, where it is confounded with low-level execution and simplified constraints. We introduce Schedule-Overcooked, a benchmark that fixes task decomposition and execution to isolate subtask ownership and dispatch timing. It pairs a four-level task-complexity suite with a five-level suite of composable agent, task, and resource constraints, and exposes the same allocation interface in TaskBench, Minecraft, and Factorio. Evaluating rule-based, LLM, and GNN-PPO allocators, we find that (i) the evaluated allocators vary substantially across settings, and near-full agent utilization often coexists with unfinished workflows; (ii) under compound temporal constraints, a GNN-PPO allocator attains low success already on its training configurations, and revealing capability information at evaluation time gives little benefit to the trained policy; and (iii) frozen policies transfer well between workload levels that share workflow structure but asymmetrically across constraint scenarios, and an Overcooked-trained policy remains operational on TaskBench, Minecraft, and Factorio, where limited adaptation further improves TaskBench success. Schedule-Overcooked provides a controlled testbed for developing allocators that coordinate time, resources, and uncertainty across concurrent tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.