Freeze, Refit, Compare: A Procedure-Level Contamination Audit for Symbolic ARC Solvers
Abstract
Symbolic solvers for the Abstraction and Reasoning Corpus (ARC) build their answers from libraries of hand-written operations, so every solution can be read. Readable code can still hide evaluation contamination: iterating on the public evaluation tasks can leave traces of them in operation bodies, validators and special cases, invisible to a check of method names for evaluation identifiers. We propose an audit of the whole executable procedure: freeze the solver, refit its trainable part on a development split of the training set, and compare success on reserved training tasks with success on the public evaluation set. Because the evaluation set is deliberately harder, an evaluation excess is a warning sign. On our own solver, the frozen pipeline's evaluation rate was 3.3 to 6.3 times its reserved-training rate on each of four splits, and commit history points to operations written during evaluation iteration as the likely source. A recount of our own provenance check found two counting errors, including evaluation tasks counted as their own held-out support because they reappear in ARC-AGI-1. With support recounted, manifests selected without evaluation outcomes and nested controls reverse the gap, as does an independently written solver; under our rule these outcomes are inconclusive, not clean passes, and for the frozen manifests the reversal is consistent with the measured difficulty shift. Under an independent-task model at reserved-training rates of 5 to 9%, the audit flags an evaluation excess with 80% probability only when the excess reaches about 9 to 11 percentage points. Inspectable is not the same as clean: contamination can hide in code that nobody calls trained, and an audit costing one training split can expose it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.