acceptodds
Under review as a conference paper at ICLR 2027

InfraMind: Infrastructure-Aware Multi-Agent Orchestration

Abstract

Multi-agent LLM systems spend additional computation on reasoning and collaboration, but shared serving queues can prevent that computation from producing answers before a deadline. This creates a coupled decision: the reasoning workflow determines the load, while changing load determines which reasoning choices remain affordable. We introduce InfraMind, a hierarchical policy that adapts workflow planning and execution to serving conditions. At arrival, a planner selects the collaboration graph using task information, load, and the time budget. Before each agent call, an executor jointly selects the model and reasoning strategy using refreshed telemetry and remaining time. A shared latency penalty connects the two policies. Our analysis characterizes why observing load matters when congestion reverses the preferred computation. Across five benchmarks on a shared five-model deployment, InfraMind improves low-load accuracy by up to percentage points over the strongest evaluated baseline. On 500 MATH questions with matched EDF scheduling and per-request deadlines of –, it increases correct, on-time answers from to . With 15 models and Azure timings replayed on the same benchmark and deadline tiers, this fraction rises from to over MasRouter. These results support adapting reasoning decisions to the serving capacity available throughout a workflow.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.