Under review as a conference paper at ICLR 2027
Candidate Selection Changes Language-Model Verifier Comparisons
Abstract
Candidate selection changes measured verifier gaps. On a separate set of 2,250 new questions, switching from Qwen-selected answers to uniform-all weighting widens the Qwen–Granite option-label AUROC gap by 0.097 [97.5% CI 0.069, 0.126], replicating a 0.100 change on an earlier 700-question shared candidate bank. An independent three-family MMLU holdout finds uniform-all AUROC advantages of 0.076–0.108 over model-chosen answers; a TruthfulQA extension tests the effect on a second benchmark. First-pass probability controls distinguish these selection effects from the incremental value of verification. Verifier comparisons require shared candidates and explicit policy weights.
open until 14 Dec 2026
est. 32% chance this paper gets accepted at ICLR 2027.
Reject 68%Accept 32%
What do you think this paper will get?
All positions stay anonymous.
Related papers
Loading the map…
Discussion (0)
Sign in to comment.