acceptodds
Under review as a conference paper at ICLR 2027

Think Without Thinking: When Empty Reasoning Reveals Answerability

Abstract

Reliable uncertainty before response generation would let a language model decide whether to answer, abstain, retrieve evidence, or spend more computation without first producing a candidate. Existing query-level observers usually read the final query token, implicitly treating observation position as fixed. We show that reasoning-model templates contain a more informative alternative: the first hidden state after the model's native empty-reasoning boundary. We call the resulting observer THInKUQ. It visits this boundary with a fixed teacher-forced suffix, decodes no tokens, and fits a lightweight correctness readout. On Qwen3-4B, THINKUQ improves subject-OOD MMLU AUROC by 9.95 points and transfers without target tuning to MMLU-Pro. Across four Qwen3 scales and tasks spanning factual recall, arithmetic, and code, the gain appears when task-relevant correctness is internally available before generation. Matched ordinary suffixes, official template variants, learned prefixes, activation patching, attention analysis, and a controlled Tagged- CoT training trajectory localize the effect to a learned post-boundary transition. In the positive regime, this transition widens correct-incorrect separation, shortens prequential codelength, and remains useful with few source labels and modest prefill overhead. These results establish observation position as a design variable for pre-answer uncertainty and provide a source-side diagnostic for recognizing when a native reasoning boundary exposes a useful correctness readout.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.