acceptodds
Under review as a conference paper at ICLR 2027

When to Overrule the Majority Vote? Precision Under Disagreement in Test-Time Reasoning

Abstract

LLM-as-a-judge, re-solving, and other correction mechanisms are often used to improve self-consistency by overruling its majority vote. These methods are usually judged by overall accuracy, which is a crude measure because a net gain can hide the right answers they break behind the wrong ones they fix. We instead measure a method's precision under disagreement, meaning how often it is right when it overrules the vote, and compare it with how often the overruled vote was right on those same questions. We study three open reasoning models on competition math, focusing on questions where their samples split. On the test contests, the correct answer is among the five most common answers for only 38% of wrong votes, so choosing among candidates has limited headroom. Unanimous same-model verification is 79.5% precise overall but right on only 17% of its overrides, while the votes it overruled were right on 50% of them. Used alone, the model's own executed programs fix about as many answers as they break, but when they all agree on a different answer, they are almost always right. Overriding the vote only in that case raises test accuracy by 1.6 points and breaks a right answer only once across all contests.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.