Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
Abstract
Test items from public benchmarks leak into pretraining corpora, and once memorized they would inflate evaluation scores. Contamination mitigation evaluation intervenes in decoding to suppress memorization and restore a contaminated model's genuine capability. However, its prevailing metric, the G-AP (Gap of Aggregate Performance), has three flaws. First, a discrete correct/incorrect readout cannot characterize per-question performance. Second, averaging before differencing lets over- and under-suppression cancel. Third, uniform per-question weighting rewards restoring the clean model's most frequent solve probabilities while neglecting the rarer ones, which differ in how hard they are to restore. To address these issues, we propose SA-PPG (Stratified Aggregate of Per-question Probability Gaps), which estimates each question's solve probability by sampling, differences it against the clean model per question, and aggregates the resulting gaps within groups defined by the clean model's solve probability. Further, on the mitigation side, since existing strategies first estimate where contamination lies and then act on that estimate, they are only as accurate as the estimate. Our strategy, RailCap, instead judges contamination during generation: whenever a sample falls back onto the greedy trajectory, the next trajectory token is capped to the runner-up, and suppression accumulates until the response distribution is sufficiently dispersed. Extensive experiments shows that SA-PPG excel G-AP in estimating the restoration of prior strategies, and RailCap attains the lowest SA-PPG in all nine settings (three models three contamination domains). Code is available at https://anonymous.4open.science/r/railcap.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.