acceptodds
Under review as a conference paper at ICLR 2027

Multi-Agent LLM Serving Is a Full-Stack Problem: A Small Step Towards Algorithm–System Co-Design

Abstract

Large language models are increasingly deployed as multi-agent systems (MAS) that solve a task by coordinating many LLM invocations through structured control flow such as branching, debate, and iterative refinement. This capability imposes a serving cost that current infrastructure is ill-equipped to absorb: algorithm researchers iterate on MAS designs rapidly and optimize for task accuracy, whereas serving systems evolve slowly and cannot be specialized to each workflow, so MAS run on general-purpose engines blind to workflow structure. We introduce a unified graph representation that casts existing and emerging MAS algorithms as typed graphs, turning opaque multi-agent executions into analyzable system workloads. On this representation we develop two families of optimization: static patterns, recurring motifs that a runtime accelerates without changing the algorithm, and dynamic strategies, small structural changes that trade accuracy for large efficiency gains. On real GPUs shared-prefix reuse cuts prefill latency by –, confidence-gated early exit reaches – speedup at – fewer tokens, and sparse expert dispatch lowers cost roughly for points of pass@1. Crucially, the prefill win does not survive end to end: decode is – of latency here, so end-to-end speedup stays at parity (–, every CI covering ), making phase-awareness a prerequisite for co-design. We intend this work as a guide for designing hardware-friendly MAS algorithms and a step toward algorithm–system co-design for LLM serving.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.