CALE: Cost-Aware Large-small Ensemble for Heterogeneous LLM Annotation Pools
Abstract
Large language models (LLMs) are increasingly used as annotators, and organizations often have access to heterogeneous pools of large and small LLMs with different costs and capabilities. We study the practical yet underexplored setting of annotating a fixed pool of unlabeled instances under a finite budget, where both labels and model reliability are initially unknown. The goal is to allocate this budget across models and instances to obtain high-quality annotations. We propose CALE, a cost-aware large–small ensemble framework that couples label aggregation with sequential acquisition. CALE models heterogeneous annotator reliability hierarchically, sharing information across similar instances while preserving instance-specific variation, and treats repeated calls to the same model as dependent rather than independent evidence. It then selects each model–instance query using complementary gains from improving the current label and learning reliability for future decisions, normalized by query cost. The resulting closed-loop interaction allows aggregation and acquisition to reinforce each other throughout the annotation process. Across five benchmarks spanning topic classification, clinical inference, entailment, ontology classification, and multi-subject question answering, annotated by heterogeneous LLM pools, CALE achieves a higher area under the accuracy–cost curve (AUBC) than every other baseline on every dataset, with the largest gains at tight budgets: it reaches 85.5% mean final accuracy at 1.5N, whereas competing systems require about 2.9N to reach the same level. Ablation and robustness studies analyze the sources of the gains and assess stability across modeling and deployment choices.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.