AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing
Abstract
Large language models (LLMs) are widely used as judges for open-ended generation because human evaluation is costly and hard to scale, yet their preferences remain imperfect proxies for human judgment. Human verification can reveal which judge decisions to trust, but it is typically affordable for only a small fraction of them. We propose AURA, an adaptive, uncertainty-aware framework that jointly refines estimates of judge–human agreement and guides human verification for pairwise LLM-as-a-judge decisions. The framework updates a probability of agreement for each comparison using a learned human-consistency scorer and local and anchor-based evidence. Sparse transport propagates reliable signals to uncertain comparisons, while uncertainty guides the allocation of additional human verification. Our synthetic and real-data evaluations show how performance varies with judge quality and verification budget, and illustrate when adaptive refinement can help allocate limited human review.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.