acceptodds
Under review as a conference paper at ICLR 2027

Recipe-Matching, Not Equivalence

Abstract

MathNet-Retrieve tests whether a retriever can find, for a math problem, a document that states the same problem. It builds those documents itself: an LLM under one fixed prompt writes each gold document and its near-miss distractors, and LLM judges filter the result. We call that procedure the benchmark's recipe, training on pairs built the same way recipe-matching, and ask how much score it buys beyond the ability the benchmark claims to test. Two models trained from the same base on the same number of rows under the same settings differ only in the training file: pairs written with the benchmark's published prompt by a different vendor's LLM and judge, or pairs a computer-algebra system verified with no LLM anywhere. The first beats the second by R@1 points on the easy tier. As far as a non-LLM paraphrase control can tell, half to two thirds of that gap comes from the training pairs being LLM-written at all: LLM rewrites under two prompts unrelated to the benchmark, with the verified model's negatives, recover and of the points, while back-translated paraphrases with the same negatives recover almost none. The remaining to points, depending on the comparison, appear only under the benchmark's own prompt, and they vanish on real duplicates no generator wrote, the same problem printed in two languages. The hard tier rewards the recipe's pair structure, a deep rewrite against a minimal-edit near-miss: LLM rewrites alone score zero on it, attaching negatives unlocks it, and every kind of negative that does so costs cross-language points against the same rewrites trained alone; the training sets that score highest on it discriminate near-misses no LLM wrote worse than plain LLM rewrites with verified negatives. MELD also moves when a model trains on pairs built its way, without losing retention; on SABER-Math the registered attack fails, and the one gain, from training on its LLM-written summaries, is small and confounded; only on MathNet-Retrieve could we pin an inversion, benchmark score up and real retention down, to one edit of a training file. We release the generator-free duplicate evaluations, the near-miss test and three trained models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.