Chorus: Learning Reliable Supervision from Noisy LLM Annotators
Abstract
Large language models (LLMs) offer a scalable alternative to human annotation, but their reliability varies across models and tasks, and their errors are often class-dependent. Such heterogeneous errors make it difficult to obtain reliable supervision for downstream learning. To address this challenge, we propose Chorus, a closed-loop framework for learning reliable supervision from noisy and heterogeneous LLM annotations. It couples supervision estimation with downstream model training and uses the resulting evidence to guide subsequent LLM annotation. In Chorus, unlabeled data is processed in sequential annotation rounds. Within each round, a downstream-aware Bayesian estimation combines LLM annotations with task-model predictions, used as instance-specific label priors, to jointly infer latent labels and annotator-specific reliability. Across rounds, Chorus identifies recurring discrepancies between individual annotations and high-confidence inferred labels, and converts them into annotator-specific feedback for the next batch. In this way, Chorus operates through two coupled loops: supervision estimation and task-model training alternate within each round, while inferred supervision is converted into feedback that guides annotation across rounds. Extensive evaluations across five datasets show that Chorus improves annotation and downstream accuracy by up to 5.29% and 6.20%, respectively. With 3–8B annotators, it approaches clean-supervision performance on multiple datasets and remains competitive with methods using 72–78B annotators, while achieving further gains with stronger annotators.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.