acceptodds
Under review as a conference paper at ICLR 2027

From Benchmark to Deployment: Efficient Evaluation of LLMs for Health under Task and Evaluator Shift

Abstract

Large language model (LLM) benchmarks evaluate models under a fixed task distribution and scoring protocol, both of which may differ from those relevant to deployment. We study how to reuse benchmark scores for deployment-specific model selection and performance estimation when scoring model responses on target tasks is resource-intensive. First, without target scores, we transport benchmark evaluations by reweighting tasks using semantic representations of prompts and rubrics, without requiring annotated shift attributes explicitly. On HealthBench, this approach can match or outperform reweighting based on known shift attributes in target-ranking recovery and mean-score estimation. Second, when a small number of target tasks are evaluated by the original benchmark evaluator, we combine their scores with benchmark scores to improve target-mean estimation and model-ranking recovery. Third, we address evaluator shift, a practical challenge when deployment evaluation relies on a human expert or a newer or more scalable LLM judge instead of the benchmark evaluator. Paired scores on a small target sample allow us to exploit the association between two evaluators' scores to improve target-mean estimation. We derive efficient influence functions for the target mean score in both the same- and different-evaluator settings. Building on these, we develop the Covariate-Adaptive Fusion Estimator (CAFE). Rather than separately estimating complex nuisance functions and plugging them into the efficient influence function, CAFE directly learns how to combine benchmark and target scores through variance minimization. Under suitable conditions, CAFE achieves semiparametric efficiency and supports asymptotically valid inference. Our experiments show that it improves target-mean estimation over existing transport and calibration approaches. Together, these results provide a principled route from static benchmark scores toward deployment-specific model selection and performance assessment under limited evaluation budgets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.