acceptodds
Under review as a conference paper at ICLR 2027

Can LLMs Improve Their Own Harness? A Controlled Study on ARC-AGI-3

Abstract

A generally intelligent system should acquire new skills efficiently without requiring humans to redesign its software scaffold for every new task. Yet language-model agents depend strongly on human-engineered harnesses that determine perception, memory, tools, recovery, and inference allocation. ARC-AGI-3 offers a concrete test of this gap: human-steered systems already solve all public levels while using fewer actions than human baselines, so effective designs are known to exist. We ask whether frontier models can build and improve such a harness autonomously. Through a fixed coding bootstrap, each model constructs a complete task harness, acts through that harness on five development games, and revises the entire codebase twice. Every version is evaluated on twenty games that never influence adaptation. We measure task performance with Relative Human Action Efficiency (RHAE), which combines level completion with action efficiency relative to human baselines, and record dollar-denominated inference cost as a practical resource measure. Across the study, models construct diverse and sometimes highly capable harnesses, but substantial feedback-driven revisions do not consistently improve held-out performance. Development gains often fail to transfer, and reductions in inference cost can come at the expense of task performance. For models evaluated in both settings, even revised harnesses fall well short of reported human-steered systems. These findings reveal a gap between autonomous harness construction and reliable adaptation: even strong task performance does not ensure that subsequent revisions generalize or preserve existing capabilities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.