acceptodds
Under review as a conference paper at ICLR 2027

ARC-Golf Bench: Benchmarking Continual Program Optimization with Code Golf Agents

Abstract

Code benchmarks typically reward correctness, leaving agents' ability to improve already-correct programs underexplored. We introduce ARC-Golf Bench, a benchmark for continual program optimization in which coding agents solve ARC-AGI grid-transformation tasks while minimizing Python program length. ARC-Golf Bench comprises 96 tasks from the NeurIPS 2025 Google Code Golf Championship, selected through item response theory-based difficulty discrimination. We evaluate four models with two agentic harnesses, tracking the shortest verified solutions relative to human performance as output-token budgets accumulate. Analyses of solution diversity and verification behavior complement these outcome trajectories. Model rankings vary across harnesses, and early progress and final quality can favor different systems. Although agents remain behind the strongest human competitors on aggregate, they collectively improve on the best known human records on 7 tasks. We further study experience transfer through a fixed library of local simplifications, conditional rewrites, and representation strategies. On ten held-out tasks, the library reduces total program length by 20.4% for one worker. For another, combining the library with an adapted harness yields a 9.4% reduction relative to library-free parallel exploration. Both comparisons show improvements on eight of ten tasks. These findings highlight the joint role of model capabilities and harness design in effective skill transfer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.