acceptodds
Under review as a conference paper at ICLR 2027

Detecting Benchmark Memorization in LLMs on Numerical Optimization Tasks

Abstract

Large language models (LLMs) give accurate answers on classic test functions. Shown samples of an unnamed function, they often return its textbook optimum to four decimals, which raises the question of whether they are inferring the answer from the data or recalling the benchmark. We propose a memorization-controlled protocol for numerical optimization on continuous test functions, with integration and multi-fidelity prediction as control tasks. Our protocol includes: (i) a recall signature, defined as the near-exact answer, which no data-driven surrogate attains on any optimization or multi-fidelity cell; (ii) a procedural generator that moves each benchmark optimization function away from its published form by adding a random smooth field, whose amplitude in units of the function's own variability sets the distance, yielding a measured recall dose-response; and (iii) a surrogate floor that distinguishes anchoring on the benchmark answer from defensible shrinkage toward it. Across sessions with seven LLMs, models returned near-exact answers on benchmark optimization problems in up to 88% of sessions, even when given fresh samples and opaque identifiers. In sessions on novel functions, they never did. Models returned the benchmark answer less often as functions moved farther from the benchmark, and how quickly this happened varied by model. GPT and Claude models returned the benchmark answer in 23–78% of sessions even when it was strictly worse than an estimate fitted to the samples shown, and this held across every surrogate floor we tested. This suggests that models recognize a benchmark when deviations fall within 1–5% of the spread in the samples. Newer Claude and GPT models relied on benchmark answers more often than their predecessors, and for Claude, higher reasoning effort increased it as well; code execution did not eliminate recall. Whether recall occurred depended on the particular samples shown as well as on the function. Therefore, accuracy on classic test functions alone does not establish numerical reasoning; the protocol offers a reusable way to evaluate numerical performance while controlling for memorization.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.