Cost- and Quality-Aware Pruning of Multi-Agent LLM Systems
Abstract
Multi-agent large language model systems often incur substantial inference costs, even when additional agents contribute little to the final answer. We study which agents can be removed from a deployed system while maintaining a specified quality requirement. We apply a novel extension of the classical random-walk model in which link traversals carry costs and rewards, using closed-form first-passage calculations to evaluate sequential execution. The process ends when the current answer is returned by a reporting manager. Traversal costs represent the accounted model-call cost, while quality increments represent changes in answer correctness whose sum equals final correctness. We first calibrate agent behavior, then evaluate candidate subsets, and construct a pruning sequence using predicted quality and cost changes. We evaluate candidate subsets using closed-form matrix calculations, thereby avoiding expensive trajectory simulations and additional model calls. Finally, we deploy the least costly candidate design whose bootstrap screening percentile meets the quality target. In controlled over-provisioned systems, the pruning method reduces accounted API cost by 73% on competition mathematics, 15% on code repair, and 10% on heterogeneous code repair. Additional mathematics experiments with Qwen, Kimi, and GLM yield cost reductions of 34.8%–73.5% at a quality target of .80. The selected designs achieve held-out accuracies at least as high as the corresponding full systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.