acceptodds
Under review as a conference paper at ICLR 2027

VeRA: Renewing Reasoning Benchmarks with Executable Specifications

Abstract

Reasoning benchmarks need fresh questions that preserve the tested computation and harder questions as models improve. We introduce VeRA, which compiles an existing problem into a reusable question template, input generator, and deterministic answer program. VeRA-E varies inputs and wording within a computational family; VeRA-H modifies the computation; VeRA-H Pro selects difficult candidates. Program checks and independent item audits validate the resulting releases. Across 16 models, VeRA-E leaves mean accuracy nearly unchanged on GSM8K (94.85% to 95.20%) and Beyond-AIME (58.34% to 57.30%), but reduces AIME-2024 accuracy from 84.46% to 70.25%. This selective drop makes computation-preserving renewal a diagnostic for benchmark-specific overfitting and possible data contamination. For modified tasks, a judge selects from up to five validated proposals per original problem. On AIME-2024-II, the selected audited release lowers accuracy from 84.91% to 58.57%. Within this budget, selection yields harder releases than the broader pools on all three sources; on AMO-Bench, the selected release's mean accuracy remains above that on the originals. Targeted repair raises usable construction yield from 75.4% to 95.1%. We release executable specifications, audited items, and validation records for repeated benchmark renewal.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.