acceptodds
Under review as a conference paper at ICLR 2027

When Can Scores Improve a Vote?

Abstract

A language model can answer a problem by generating several solutions and choosing the most common final answer. We characterize exactly when scores make voting errors decrease faster as more solutions are sampled. The key information is how scores distinguish the correct answer from its leading wrong rivals. We study independent samples for one fixed problem with finitely many possible answers, where errors occur but the correct answer is more likely than any individual wrong answer. Some fixed rule with positive, bounded score weights makes errors decrease faster than ordinary voting at an exponential rate if and only if the correct answer’s score distribution cannot be reproduced by mixing those of the most frequent wrong answers. Two perfectly calibrated examples (whose scores match correctness frequencies) have identical score–correctness statistics and answer frequencies, yet differ in this capacity for improvement. Complementing this large-sample characterization, we identify the weights that preserve the long-run winner under calibration and give conditions for improving count ties safely at finite budgets. On a shared cohort of 252 competition-mathematics problems, score weighting improves eight-candidate accuracy by 2.78 percentage points for Math-7B (95% interval [1.06, 4.63]) and 2.91 for Phi-4 ([1.19, 4.76]). Count ties account for most of both gains. Further experiments evaluate scoring cost, transfer, and score-based diagnostics.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.