acceptodds
Under review as a conference paper at ICLR 2027

Repetition Is Not Accuracy in Large Language Model Evaluation

Abstract

Repeated queries to a large language model often produce differently worded answers with the same meaning, and evaluators use this semantic agreement as a black-box reliability signal. We ask whether agreement can support an accuracy claim when labeled source prompts and unlabeled deployment prompts come from different populations. A model can repeatedly return the same wrong answer, so agreement measures concentration rather than truth. We capture this distinction with a promptwise agreement bias: the amount by which agreement overstates or understates accuracy. Under an oracle law for labeled source evaluation, the worst-case ambiguity is exactly the deployment mass outside source coverage. When source information is reduced to mean agreement bias and a justified promptwise range, we derive the full compatible risk interval in the corresponding summary model. Conditional shift bounds and finite-sample intervals then separate the effects of sampling new prompts from sampling additional responses to a fixed prompt. Matching constructions show when more responses help, when broader prompt coverage matters, and when target labels or external assumptions remain unavoidable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.