Beyond Benchmark Accuracy: Stress-Testing Reasoning Claims in Large Language Models
Abstract
Large language models, and reasoning models in particular, score highly on mathematical, logical and coding benchmarks, and those scores are increasingly read as evidence of general reasoning ability. Accuracy on a fixed item set does not establish that the underlying competence survives controlled variation in problem depth, surface form or evaluation condition. We stress-test reasoning claims along four dimensions — complexity generalization, perturbation robustness, process faithfulness and familiarity–novelty sensitivity — and add a cross-dimensional analysis of how the first two interact. Fourteen contemporary configurations over ten model checkpoints, including four reasoning-disabled counterparts of listed checkpoints, are evaluated on two generated task families with a depth ladder reaching , on five benchmark reference points scored as integer item counts, and under transformations separated by whether they preserve the correct answer. Benchmark accuracy and stress behaviour order the roster similarly overall (Spearman ) and come apart where a benchmark discriminates least: configurations within 5\ percentage points of one another on MATH-500 differ by up to 23.8\ points in their answer-preserving robustness gap, and the two orderings disagree on 8 of 91 configuration pairs. The incremental information is local to near-ties rather than a general decoupling, and it is smaller among distinct checkpoints than across all configurations. The perturbation penalty compounds with depth rather than adding to it — the fitted interaction is negative for every configuration, and for the strongest the gap widens from 0.8 points at to 11.3 at . The direction holds under depth entered as a factor, as a spline and segmented at the changepoint, under matched input length, and under inference clustered on seed templates rather than items: 184 of 192 intervals lie below zero, the exceptions being the two strongest configurations under the widest specifications. Threshold-like degradation is identified for ten configurations with the changepoint tracking capability, right-censored for the two strongest and unidentifiable for the two smallest, so a single ladder cannot serve the whole capability range. These are behavioural results about generalization under stated conditions, not findings about internal computation. Within that scope they indicate that benchmark accuracy and robust generalization should be reported as distinct empirical properties, and that a robustness estimate obtained at one difficulty level understates fragility at greater depth.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.