Two Calls, Two Moments, and the Vote-Accuracy Curve of Repeated LLM Inference
Abstract
Repeated sampling can improve LLM accuracy, and quantifying the gains from additional calls is essential for allocating test-time compute. We study binary decisions, a fundamental setting where repeated answers to the same question are aggregated by majority vote. We show that two independently sampled responses per example in a validation set with known answers constrain the latent distribution of example-level success probabilities to a class consistent with the paired outcomes. Optimizing over this class yields sharp accuracy and gain bounds at every finite voting budget. A shared large-sample confidence region accounts for validation uncertainty. We also obtain sharp infinite-vote bounds and moment-matched forecasts of the vote-accuracy curve. On QNLI and QQP, our method successfully distinguishes settings with voting gains above a target margin from those with little room for improvement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.