Compiled Worlds: What Adaptation Costs on Transformed ARC-AGI-3 Games
Abstract
Completion on familiar games does not show how an agent adapts when a game changes, or what the adaptation costs. We introduce Compiled Worlds, a protocol that edits official ARC-AGI-3 game programs into paired variants (colour permutations, action relabellings, resampled layouts and rule mutations), serves them through the unmodified toolkit, and attaches executable solution witnesses to layouts and mutations. We evaluate four harnesses with a common backbone, GLM-5.3, and selected configurations with three further backbones. Because every variant has an original, the protocol separates what a harness costs on a familiar level from what a change adds. On the original first levels of the five games with layout variants, Tycho with GLM-5.3 already uses nine times the calls of a compile-by-intervention harness (CbI) with the same backbone. Measured against its own originals, Tycho pays mainly for new layouts, 3.3 times its original first-level calls and more on all eight layouts it completes, whereas CbI's premium is at most 1.4 on every tier; each fails five of thirteen layouts. First-level rule mutations cost about as much as the original levels, and each model-building configuration evaluated on them with an API backbone solves at least four of five. Across 33 paired first runs, colour permutation does not change completion measurably. Finally, ten of CbI's twelve zero-completion first runs with GLM-5.3 reach full agreement with the logged transitions, so trace agreement does not ensure completion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.