When Should LLM Agents Talk? Measuring Coordination Value and Auditing Learned Routers
Abstract
After a language-model agent produces an initial answer, the system must decide whether further coordination is worth the inference cost. Even if an initial answer is assessed as incorrect, coordination may not improve it enough to justify its cost. We define the policy-relative Causal Value of Coordination (CVC) as the expected difference in cost-adjusted utility between a coordination action and a feasible alternative. Both start from the same history and follow the same fixed continuation policy. With the incumbent policy as the reference, a performance-difference identity links exact CVC values to end-to-end policy improvement. We estimate CVC through paired counterfactual replay and train ridge and neural predictors to construct candidate routing policies. Frozen candidates are assessed through full-policy comparisons and audits. Our main evaluation uses 2,000 MMLU-Pro questions across two backbones, complemented by supplementary GSM8K and MMLU-Pro experiments. Confidence predicts initial errors, but the prespecified router comparisons do not establish a utility gain. A diagnostic with access to initial correctness yields larger gains, highlighting the gap between error prediction and effective routing. A fixed early-stopping policy reduces voting cost without changing answers; observable-failure recovery yields modest secondary gains. Learned-policy lower bounds remain negative, and no candidate passes the supplementary independent audits. These findings highlight that identifying beneficial coordination, learning effective routing policies, and certifying end-to-end improvements are distinct challenges under the tested conditions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.