acceptodds
Under review as a conference paper at ICLR 2027

Continual-ARC: Measuring Skill Acquisition over a Working Lifetime

Abstract

Agent harnesses have emerged as a promising mechanism for continual learning, alongside test-time parametric learning. However, the current benchmarks that measure such systems have several limitations: they generally focus on a single applied domain, they serve each instance only once, and most assume a single family of learner, so memory-based and weight-based systems are rarely compared on the same scale. We propose Continual-ARC, a lifelong stream built from the three ARC-AGI editions, whose tasks need no domain knowledge and which the models under test solve only in part. The same tasks return after dormant gaps with fresh instances, while new tasks keep arriving. On each instance, the learner may pay for the demonstrations or submit an answer, with retries permitted on an incorrect submission. Each instance is scored based on the share of this feedback budget (demonstrations and failed attempts) consumed. Since a learner may keep anything between instances, memory-based and weight-based systems are scored on the same scale. We run three agent harnesses with frozen weights, each with and without persistence, comparing them with two prompting floors and a program-induction method. We also define a second regime in which one system learns from several users simultaneously. Across our experiments, we show that persistence in the harness does generally save on budget, but only modestly. None of them perform measurably better than a floor that simply caches each task's demonstrations. In contrast, a simple method that asks the LLM for one program per task and keeps following demo reproduction outperforms all harnesses by a wide margin. This opens the door to future harness design improvements, and to weight-based continual learning approaches that could teach models to effectively drive harnesses in new environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.