acceptodds
Under review as a conference paper at ICLR 2027

Paying for Better Data in Prediction-Powered Testing

Abstract

Automated judgments offer a cheaper alternative to expert labels for evaluating large language models, but their accuracy may depend on a provider's unobserved effort. We study whether paying for better judgments can reduce testing costs. Our design combines random audits with bonuses for agreement with experts and chooses the contract and sample size together. We show that the robust likelihood factor maximizing expected log evidence equals a prediction-powered betting factor with jointly optimized predictions and bet size. The test controls false positives in finite samples under arbitrary reporting and attains the optimal leading sample size under the induced review model. With uncertain review costs and accuracy, we characterize contracts and sample sizes ensuring participation, high effort, and prescribed power throughout specified ranges. A finite-sample cost criterion compares candidate designs with direct gold acquisition. Simulations identify when incentives reduce certified costs, and comparisons with prediction-powered baselines clarify the role of inference efficiency. An LLMBar study using observed judgments and assumed costs illustrates how a weaker accuracy gain can favor direct gold acquisition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.