acceptodds
Under review as a conference paper at ICLR 2027

Depth Over Breadth: Energy-Efficient Scaling of Multi-Agent LLM Systems

Abstract

Deploying a multi-agent LLM system (MAS) means choosing how far to scale three axes: per-agent tool-calling depth , agent count , and communication rounds . Each choice buys accuracy and burns energy, yet prior scaling studies price compute in tokens or dollars on closed API models where energy cannot be measured, and tokens misprice it: a decode token costs 335 a prefill token at short context under our setup. We measure hardware-level GPU energy at every LLM call across four topologies, four agentic benchmarks, and three locally served models (9B to 35B, the larger two on SWE-bench Lite only). We find that test-time scaling axes are not interchangeable: tool-calling depth () converts energy into accuracy most efficiently; agent breadth () helps less and costs energy linear in the number of agents; and communication rounds () buy no measurable accuracy once per-agent iterative depth is held constant (for homogeneous agents), while having steep energy cost. Based on our findings of this depth-over-breadth ordering, we develop a label-free configuration search-and-deploy policy: binary-search on a single agent to the deepest budget at which agents still utilize at least half of it, then deploy the target MAS topology at that budget with three agents. Against a two-phase depth-then-breadth scaling policy, this cuts search-and-deploy energy by 59% on average while giving up at most half a point of accuracy on three of four benchmarks; and against running the same label-free rule on the MAS topology instead of a single agent, our policy saves 51% at comparable accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.