acceptodds
Under review as a conference paper at ICLR 2027

Cost- and Quality-Aware Pruning of Multi-Agent LLM Systems

Abstract

Multi-agent large language model systems often incur substantial inference costs, even when additional agents contribute little to the final answer. We study which agents can be removed from a deployed system while maintaining a specified quality requirement. We apply a novel extension of the classical random-walk model in which link traversals carry costs and rewards, using closed-form first-passage calculations to evaluate sequential execution. The process ends when the current answer is returned by a reporting manager. Traversal costs represent the accounted model-call cost, while quality increments represent changes in answer correctness whose sum equals final correctness. We first calibrate agent behavior, then evaluate candidate subsets, and construct a pruning sequence using predicted quality and cost changes. We evaluate candidate subsets using closed-form matrix calculations, thereby avoiding expensive trajectory simulations and additional model calls. Finally, we deploy the least costly candidate design whose bootstrap screening percentile meets the quality target. In controlled over-provisioned systems, the pruning method reduces accounted API cost by 73% on competition mathematics, 15% on code repair, and 10% on heterogeneous code repair. Additional mathematics experiments with Qwen, Kimi, and GLM yield cost reductions of 34.8%–73.5% at a quality target of .80. The selected designs achieve held-out accuracies at least as high as the corresponding full systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.