Answer Extraction Changes Conclusions in Controlled Mathematical Reasoning Tests
Abstract
A mathematical answer score combines the response a model produces with the rule used to read that response. We separate these two components in a controlled study of six mathematical template families. The dataset contains 240 programmatically checked problems, each presented with and without a task-specific reasoning cue. Qwen2.5-7B-Instruct and Llama-3.1-8B-Instruct produce 960 responses under a common decoding protocol. A strict final-answer rule assigns correct scores to 217 of 480 Qwen responses and only 2 of 480 Llama responses. Yet the rule fails to extract an answer from 477 Llama responses, including texts that state a numerical conclusion in ordinary language. A post hoc, label-blind terminal-answer sensitivity recovers 242 additional Llama predictions, raising its correct count to 135 without changing any generated text. Under this sensitivity, the mean cue effect is percentage points for Qwen and points for Llama, with descriptive family-level intervals spanning zero. Individual family effects vary substantially. These results do not establish either a universal cue benefit or a model ranking in mathematical reasoning. They show why a benchmark must distinguish answer-format compliance, extractable correctness, and the effect of a reasoning intervention. We release paired prompts, exact labels, complete responses, and both scoring rules so that these quantities can be recomputed separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.