acceptodds
Under review as a conference paper at ICLR 2027

Compute-Optimal Stopping for Multi-LLM Deliberation

Abstract

Multi-LLM systems can spend substantially different amounts of inference compute on the same query, making adaptive stopping essential for balancing answer quality and computational cost. We present Dwell, a cost-sensitive stopping framework for selecting among fixed-depth multi-LLM deliberation policies. Under diminishing marginal accuracy per unit cost on the efficient frontier, we show that the cheapest utility-maximizing endpoint follows a threshold rule and that dominated depths can be excluded as stopping endpoints. In a controlled five-seed simulation, the selected depth-four endpoint achieves 0.950 ± 0.011 accuracy at a normalized cost of 7.1, compared with 0.964 ± 0.010 accuracy at a cost of 19.2 for mixture-of-agents. The resulting operating point is 2.7× cheaper while sacrificing only 1.4 accuracy points. A lightweight confidence model obtains an expected calibration error of 0.102, and our conditional headroom bound connects calibrated confidence to the value of further deliberation. A held-out 60-item ARC-Challenge study with three API-served workers further characterizes how the preferred endpoint changes with the cost coefficient λ. Together, these results establish a principled mechanism for selecting deliberation depth under explicit accuracy–cost trade-offs while separating controlled evidence from the preliminary live-model evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.