acceptodds
Under review as a conference paper at ICLR 2027

Finding Answers in Thought Matters: Revisiting Evaluation on Large Language Models with Reasoning

Abstract

Evaluating generative models, such as large language models (LLMs), commonly involves multiple-choice question-answering tasks where the final answer is selected based on the probabilities of the answer choices. For reasoning models, however, the method of answer extraction plays a critical role. In this paper, we show that the benchmark performance of reasoning models and their extracted-answer distributions are highly sensitive to the answer-extraction algorithm employed. To mitigate this sensitivity, we propose a simple framework: Answer Regeneration. The method uses an additional inference step, providing the evaluated model with its original input and previous output, followed by the prefix "Answer:". The final answer is then selected using option probabilities or extracted from the regenerated output. We show that this reduces reliance on handcrafted extraction rules and achieves competitive performance with specialized extractors. Furthermore, it aligns more closely with human interpretation on sampled disagreements and provide reliable performance on incomplete reasoning responses. This framework also extends to mathematical reasoning and open-ended question answering. Our analysis and framework offer a practical approach to more reproducible evaluation of reasoning models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.