acceptodds
Under review as a conference paper at ICLR 2027

Candidate Selection Changes Language-Model Verifier Comparisons

Abstract

Candidate selection changes measured verifier gaps. On a separate set of 2,250 new questions, switching from Qwen-selected answers to uniform-all weighting widens the Qwen–Granite option-label AUROC gap by 0.097 [97.5% CI 0.069, 0.126], replicating a 0.100 change on an earlier 700-question shared candidate bank. An independent three-family MMLU holdout finds uniform-all AUROC advantages of 0.076–0.108 over model-chosen answers; a TruthfulQA extension tests the effect on a second benchmark. First-pass probability controls distinguish these selection effects from the incremental value of verification. Verifier comparisons require shared candidates and explicit policy weights.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.