HOW MUCH INDEPENDENT INFORMATION DOES AN LLM COMMITTEE PROVIDE?
Abstract
Adding agents to an LLM committee can improve accuracy without providing proportionally more independent information. Member accuracy, error dependence, and aggregation quality change together, obscuring what a larger or more communicative committee contributes. We study this problem with a committee information audit. End-task accuracy alone cannot tell whether a gain comes from better individual answers, less redundant errors, or better selection of existing answers. We therefore separate three quantities on frozen output records within each protocol arm: dependence among member outputs, complementarity available in that arm’s candidate answers, and the gain recovered by an aggregator. This separation distinguishes information that is present in the committee from information that a deployable policy can actually use. The audit combines explicit evidence allocation, matched communication contrasts, and held-out evaluation of fixed-candidate selectors. Across controlled domains, held-out template families, and two external benchmarks, variance-equivalent effective size shows saturation. In a paired Qwen3-4B full-topology control, the observed decay is steeper than a fixed-correlation null. Five related diagnostics show the same directional pattern on the shared output matrix. Conditioning on input structure explains much of the unconditional dependence, while quality-matched model diversity does not consistently reduce it. In the original output-budget-matched Qwen comparisons, peer discussion increased plurality accuracy over self-revision by 5.4 percentage points on HotpotQA and 7.4 on MuSiQue, while increasing error correlation and reducing candidate-oracle coverage. In the stricter input-token-matched Qwen validation, the peer–self plurality changes were +0.4 pp (95% CI [−2.4, 3.0]) on HotpotQA and +4.0 pp (95% CI [2.0, 6.0]) on MuSiQue; fixed-candidate oracle coverage fell by 15.2 and 10.6 pp, respectively. Grouped validation further shows that weighting and learned selection recover useful but uneven fractions of the available complementarity. Thus, communication can improve an endpoint score while increasing error correlation and reducing the fraction of questions with at least one correct candidate. Committee evaluation should therefore report dependence, candidate coverage, and realized gain together rather than use accuracy as a proxy for all three.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.