Posterior-Level Aggregation for Multi-LLM Multiple-Choice Decision Making: A Blackwell Perspective
Abstract
The rapid development of large language models (LLMs) has motivated research on answer aggregation in multi-LLM systems, where multiple models contribute predictions toward a shared decision. Existing aggregation approaches, such as voting and debate, are often developed without a common formal account of the information retained by their outputs. In this paper, we analyse multi-LLM answer aggregation (MLAA) using Blackwell’s informativeness framework. Within the Blackwell information-structure abstraction, we show that voting and debate induce information structures that are no more informative than the joint private information of all agents. More generally, this result establishes joint private information as a protocol-independent upper bound, while the corresponding ideal decision is based on the posterior conditioned on the joint private information. Motivated by this analysis, we develop a practical framework that estimates option-level distributions from sampled responses and combines them through generalised Bayesian aggregation. The framework is instantiated using product-of-experts and logarithmic opinion-pooling rules. Experiments on six multiple-choice QA benchmarks show that these posterior-level aggregation methods are competitive with or achieve higher accuracy than other multi-LLM voting and debate methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.