Robust Reasoning Benchmark
Abstract
While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving ability remains sensitive to con- text and textual formatting. We introduce the Robust Reasoning Benchmark (RRB), a pipeline of 13 deterministic textual perturbations applied to AIME 2024 and AIME 2025. Evaluating 8 state-of-the-art models, we find that frontier models are largely resilient, with the notable exception of Claude 4.6 Opus, which categori- cally refuses many transformed prompts, while open-weights reasoning models degrade by up to 53% on average and up to 100% on individual perturbations. By prompting models to solve several unperturbed problems in sequence, we identify Intra-Query Attention Dilution: open-weights models from 7B to 120B parameters lose accuracy on later problems. At matched preceding token length, the same mod- els perform better when earlier reasoning is replaced with neutral text. Per-layer attention traces show persistent attention mass on the earlier problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.