BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
Abstract
AI judges are increasingly used to reduce the cost of human evaluation. While they provide a scalable and inexpensive alternative, their outputs can be biased relative to human preferences and highly item-dependent, with substantial variation across judges, tasks, and domains. When practitioners rely on uncalibrated AI evaluations for model ranking, item scoring, or population-level quality reporting, systematic bias can propagate directly into downstream decisions and make the conclusions unreliable. We propose BACON, a four-stage pipeline that combines budgeted human calibration with multiple AI-judge outputs to obtain more accurate annotations. BACON first constructs full-coverage auxiliary features for every item in the evaluation pool, including multi-judge scores, token-level uncertainty statistics, and contextual embeddings. It then obtains human labels for a small sampled subset and trains a cross-fitted outcome model to produce calibrated item-level surrogate predictions. These predictions support two downstream modes: the first is population-level estimation of summary metrics, such as means, quantiles, or other estimands, using an augmented estimating-equation estimator with valid confidence intervals; and the second is individual-level surrogate scoring for item scoring and ranking. Throughout, BACON treats AI judges as auxiliary measurements rather than ground truth: human labels provide the calibration anchor, while AI-derived signals serve as surrogate predictions and improve efficiency. We validate BACON on a variety of tasks and domains. Across settings and labeling budgets, the cross-fitted outcome model improves predictive accuracy and ranking consistency, while the estimating-equation estimator reduces bias and variance relative to raw AI outputs and purely human-label-based methods. These results suggest that BACON provides a practical framework for scalable, statistically grounded evaluation with limited human annotation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.