Knowing Is Not Choosing: What Explicit Verification Adds Beyond Generative Preference
Abstract
Generating a correct answer does not ensure that a language model will select it. We track factual competence through three stages: whether a correct candidate is generated, how candidates are ranked, and which answer is ultimately returned. Pre-generation readouts predict factual recall beyond exposure and identity features across three model families and, without refitting, which questions sampling will cover; however, they add little about whether an available correct answer is ultimately returned. Explicit verification with improves within-question ranking over mean log-likelihood in Gemma, Qwen3, and Llama, raising pair-weighted AUROC by –. In a prospectively specified Gemma cohort, verification improves plurality accuracy by points under a frozen stricter criterion and by points under blinded semantic labels. Against chat-template likelihood, a stronger generative baseline, the gain remains points ( CI ), concentrated in relations with strong common-answer priors and dependent on the entity. In two 27B Qwen models, verification exceeds chat-template likelihood by – AUROC; masking the entity removes this ranking advantage, while plurality gains remain at most points. The measured gain also depends on the correctness criterion: recall-oriented reference matching credits option lists favored by likelihood and reduces the plurality gain over mean log-likelihood to points, whereas human labels give .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.