acceptodds
Under review as a conference paper at ICLR 2027

Can More Samples Fix the Vote? Anytime-Valid Audits of LLM Self-Consistency

Abstract

Self-consistency selects the most frequent answer among repeated language model responses without a verifier. More samples stabilize voting but leave the response distribution unchanged, so the winner need not be acceptable. We propose Response Evaluation with Sequential Outcome-Level Verification via E-processes (RESOLVE), which audits whether population plurality favors PASS or FAIL under a fixed evaluator, leaving generation and selection unchanged. For a fixed policy, deterministic evaluator and i.i.d. responses, false reports remain controlled during continuous inspection. We bound response cost for an uncapped idealization and show when identifying a label can require less evidence than identifying an exact answer. Experiments distinguish answer availability, evaluator acceptance and plurality selection. An unfavorable diagnosis can motivate policy revision, which requires a fresh audit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.