I Know That I Know Nothing: Narrowing the Knowledge-Output Gap in Large Language Models through Early High-Entropy Token Intervention
Abstract
Large language models can encode task-relevant information in their hidden states while failing to express it reliably during generation. We show that, in many incorrect cases, signals predictive of later failure are already present before decoding begins, yet these internal judgements do not consistently constrain the realised output. We identify structural compression in autoregressive decoding as a mechanism contributing to this knowledge-output gap: as generation proceeds, judgement-relevant information becomes progressively compressed along causal attention paths, weakening its influence on the final decision. To mitigate this effect, we propose a minimal test-time intervention that selectively targets a small number of early high-entropy tokens, where uncertainty and path selection concentrate. The intervention delays premature collapse into erroneous reasoning trajectories while leaving low-uncertainty decoding largely unchanged. Across benchmarks covering general knowledge, programming, mathematical reasoning, scientific reasoning, and multi-hop question answering, this strategy improves reasoning performance by 3.19 to 7.29 points and, in the inference-cost analysis, reduces generated-token usage by 25.4% to 38.9% while maintaining the highest accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.