acceptodds
Under review as a conference paper at ICLR 2027

An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems

Abstract

Frontier large language models (LLMs) now reach near-ceiling accuracy on standard mathematical-reasoning benchmarks and gold-medal-level performance at the International Mathematical Olympiad. As these benchmarks saturate and their items leak into training data, a high score no longer shows whether a model reasons robustly, nor which component of its reasoning fails when it does not. To evaluate reasoning in a way that remains informative for such models and makes its failures diagnosable, we propose a novel evaluation methodology called GAP (Generalisation-and-Perturbation), where the key idea is to automatically generate many mathematically-equivalent variants of an existing mathematics problem at scale, using two disjoint interpretable transformations: 1) surface renames, which probe the binding between identifiers and latent variable roles, and 2) kernel rewrites, which probe whether the high-level proof plan survives a change of mathematical setting. Compared with existing benchmarks, a benchmark created using GAP has two important benefits: 1) Since the new variants generated are unlikely "seen" by an LLM, a GAP benchmark effectively addresses the data leakage problem. 2) A GAP benchmark enables diagnosis of reasoning failures based on the performance of an LLM over a spectrum of variants of mathematics problems designed with different transformations, each corresponding to a meaningful hypothesis about the cause of failure. We instantiate GAP on all of the 1,051 William Lowell Putnam Competition problems from 1938 to 2024, adding 5,255 unseen transformed instances to form PutnamGAP, a 6,306-item competition-level mathematics corpus, resulting in the first public machine-readable dataset built from the full Putnam archive. Using PutnamGAP, we evaluated 18 commercial and open-source models spanning different sizes and providers, and found that accuracy drops across all models and variant families, with the largest degradations under kernel rewrites. This gap does not close with model strength, suggesting that the dominant remaining weakness of even the strongest models is carrying a proof plan into a changed setting, not handling surface changes. Further diagnostic analysis enables more detailed understanding of the causes of failure and generates potentially useful insights for future research on improving the reasoning capacity of LLMs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.