Query–Key Ensembles for Reasoning Verification
Abstract
Large language models typically generate reasoning chains that are then judged from the outside, by solvers, reranking heuristics, or trained reward models. We ask whether the model has already formed a judgment of its own. Given a generated solution, we read which of two explicit verdicts, correct or erroneous, the model’s attention heads prefer, using raw query–key (QK) inner products taken before Q/K normalization, rotary position encoding, and the softmax. Building on QK readouts for multiple-choice selection and logical consistency, we turn these head-wise preferences into a verifier of complete chain-of-thought (CoT) solutions without updating the model. A calibration-free rule, CF-QK, weights each head by its agreement across read positions and its responsiveness across an unlabeled batch; a linear readout, L2-QK, needs only 50 labeled solutions. Across Qwen3 (8B–32B) and DeepSeek-R1-Distill-Llama-8B on MATH-500 and HLE-1/4, these ensembles outperform the model’s own verdict logits, solution likelihood, and prompted selfverification in most settings. Without labels, CF-QK raises balanced accuracy over verdict logits by up to 18 points. Across ten 50-label calibration splits, L2-QK gains 4–16 points over calibrated logits on every model and dataset. The fitted rules transfer without target labels to GSM8K and SVAMP, and each verdict costs one forward pass. Matched edits to solution content and activation patching show that the scores respond to the reasoning itself and that the selected heads shape the model’s verdict. Complementary experiments study multiple-choice selection with or without a “think-first” CoT preface.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.