acceptodds
Under review as a conference paper at ICLR 2027

Can Correctness Readouts Choose the Right Answer? A Question-Level Audit of Compact KV-Cache Signals

Abstract

Can a correctness readout that ranks responses well in aggregate identify which answer to choose within a question? We study this question with KV24, a compact 24-feature summary of late-layer key–value cache statistics, paired with a lightweight correctness predictor; the language model remains frozen. We read these features before the dedicated final-answer field, although reasoning may already contain the answer. On held-out GSM8K, assigning every Qwen response its question's mean score removes all within-question ordering. The area under the receiver operating characteristic curve (AUC), computed over responses pooled across questions, changes from .9166 to .9104. Within-question ranking evidence is positive for Llama but inconclusive for Qwen. Hidden-state and likelihood controls match or outperform KV24 under several comparisons. Without target refitting, the joint Qwen-Llama ranking test passes narrowly on OpenBookQA but not on a synthetic cross-domain suite. Changing the source-validation selection criterion yields no consistent answer-choice benefit across models, and same-pool majority voting matches or exceeds the selected readouts. Together, these results separate boundary readability from transport and answer-choice utility.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.