On the Calibration of Crowdsourcing Aggregators
Abstract
Calibration has been extensively studied for classifiers, but remains largely unexplored for crowdsourcing aggregation methods, which are becoming increasingly popular for making probabilistic decisions. To the best of our knowledge, we provide the first theoretical and empirical study of calibration in label aggregation. To our great surprise, we find that widely used aggregation methods can be severely miscalibrated. We prove that the Bayesian aggregator is perfectly calibrated under its data-generating model. We also show that exact calibration of weighted majority votes requires highly restrictive conditions, which helps explain why Majority Vote is generally miscalibrated in realistic crowds. Empirically, we study calibration across a large collection of human crowdsourcing datasets and LLM-based crowds. We observe, in both cases, systematic miscalibration for aggregator models such as Majority Vote, Dawid–Skene, GLAD, and CrowdFM. Finally, we study post-hoc calibration in crowdsourcing. We propose an aggregation of aggregators approach: for any aggregation function , we introduce a transformation that first combines its predictive distribution with those of complementary (under- and overconfident) aggregators, before applying Temperature Scaling. This simple and interpretable transformation makes Temperature Scaling more effective than when applied directly to , improves calibration, especially in the high-redundancy regime, where there are many annotations per item, while preserving or even improving accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.