Discarded but Not Forgotten: Attending to Unchosen Tokens for LLM Reasoning
Abstract
In chain-of-thought (CoT) reasoning with large language models (LLMs), each decoding step selects one token to extend the generated reasoning trajectory, while the remaining unchosen tokens are not retained in the generated context. As a result, the states the model would compute by processing these unchosen tokens are never produced and remain inaccessible to later decoding steps, even though they may still contain information useful for subsequent reasoning. To retain this information without generating additional reasoning trajectories, we propose Unchosen Token Memory (UTM), a training-free, inference-time decoding framework that preserves the KV states of top-ranked unchosen tokens as candidate memory. UTM makes these candidate states accessible to subsequent decoding steps through attention, calibrates their influence according to their relative probabilities at the steps where they are produced, and adaptively determines their retention period based on decoding uncertainty. In this way, UTM maintains a single generated reasoning trajectory while allowing subsequent decoding steps to selectively access previously discarded candidate states. Extensive experiments on mathematical reasoning, knowledge-intensive reasoning, and code generation benchmarks further show that UTM achieves higher performance than standard CoT reasoning and recent inference-time reasoning methods across multiple model scales. These results indicate that discarded candidate states provide a useful source of auxiliary information for improving LLM reasoning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.