acceptodds
Under review as a conference paper at ICLR 2027

TRUST: LLM Evaluation with Imperfect Judges via Reliability Transfer under Sparse Target Supervision

Abstract

Modern large-scale LLM evaluation increasingly relies on imperfect LLM judges, yet reliable estimation of model performance remains difficult when human annotations are available for only a small fraction of target examples. Current approaches either treat the judge as a black-box predictor or model the judge's error behavior, but the former suffers from bias whereas the latter is ineffective in low-supervision regimes. We address these limitations by studying how learnt judge behavior on auxiliary benchmarks can be transferred to a target benchmark to improve LLM evaluations. Concretely, we introduce TRUST, a two-stage framework that 1) aligns source and target examples in semantic-confidence space and filters poorly matched source samples, and 2) transfers judge behavior in each semantic-confidence region through an adaptively weighted likelihood to aid model performance evaluation. Our analysis characterizes how judge reliability transfer improves model evaluation, showing it depends on the tradeoff between variance reduction from additional auxiliary benchmarks supervision and bias from residual source benchmarks–target benchmark judge behavior mismatch. Across diverse LLM evaluation benchmarks, TRUST substantially reduces LLM failure-rate estimation error by up to approximately 20% relative to target-only and prediction-powered inference (PPI) baselines, with the largest gains experienced when target human supervision is extremely limited.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.