VeRA: Renewing Reasoning Benchmarks with Executable Specifications
Abstract
Reasoning benchmarks need fresh questions that preserve the tested computation and harder questions as models improve. We introduce VeRA, which compiles an existing problem into a reusable question template, input generator, and deterministic answer program. VeRA-E varies inputs and wording within a computational family; VeRA-H modifies the computation; VeRA-H Pro selects difficult candidates. Program checks and independent item audits validate the resulting releases. Across 16 models, VeRA-E leaves mean accuracy nearly unchanged on GSM8K (94.85% to 95.20%) and Beyond-AIME (58.34% to 57.30%), but reduces AIME-2024 accuracy from 84.46% to 70.25%. This selective drop makes computation-preserving renewal a diagnostic for benchmark-specific overfitting and possible data contamination. For modified tasks, a judge selects from up to five validated proposals per original problem. On AIME-2024-II, the selected audited release lowers accuracy from 84.91% to 58.57%. Within this budget, selection yields harder releases than the broader pools on all three sources; on AMO-Bench, the selected release's mean accuracy remains above that on the originals. Targeted repair raises usable construction yield from 75.4% to 95.1%. We release executable specifications, audited items, and validation records for repeated benchmark renewal.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.