Illusions of Reasoning Fragility: How Evaluation Artifacts Inflate Apparent LLM Structural Sensitivity
Abstract
The new investigations reveal that the potential cognitive power of large language models is compromised even by a slight modification in the prompt given. The present paper presents an instance of the GSM8K performance, which equals approximately 15%. The issue at the heart of such failure is two miscalculations in the evaluation rather than deficiency in the model’s abilities. The first mistake is the additional padding in the decoding process; the second mistake emerges once the instructions are given in the prompt after the question mark and inherent double question bias. Symmetric Prompt format is suggested to eliminate the above-mentioned problems. Testing the Qwen2.5-7B-Instruct model demonstrates the restoration of more than 80% accuracy, which indicates that the problem of the illusion of reasoning loss has been resolved (echo rate drops to 0.0%). Interestingly, Llama-3.1-8B-Instruct proved largely immune to the initial artifact, maintaining high accuracy throughout. Logistic regression allows us to separate the effects of structure from those of length when testing the Qwen model (p = 0.405), while for Llama the found β value is equal to 1.09 at p = 0.004. The third model, Mistral-7B, fails under the adjusted conditions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.