Active Exploration, Rule Discovery, and Teaching in ARC
Abstract
Abstract reasoning benchmarks such as the Abstraction and Reasoning Corpus (ARC) typically provide models with a small set of demonstrations and evaluate whether they can infer the underlying transformation rule. However, the selection of these demonstrations and its role in successful inferences has received comparatively little attention. In this work, we instantiate interactive environments for ARC tasks using task-specific generator and verifier programs, allowing us to study both active discovery and teaching. We introduce ActiveARC, a framework in which models can construct input grids, observe their transformations through an oracle, and then solve held-out instances. We further ask whether models that already know the underlying rule can select demonstrations that effectively teach it to another model. Across four ARC-style datasets and five model families, requiring models to actively select their own examples substantially reduces performance relative to randomly sampled examples under matched query budgets. Accuracy drops by 12.8 percentage points for official items on average and 13.6 points for freshly sampled items, with active selection underperforming random sampling in all evaluated settings. We find that models systematically construct smaller experiments than both human-authored training examples and random samples from the task generator, and in most settings less diverse ones. Surprisingly, human-authored demonstrations often provide only a modest advantage over random generator samples, especially on freshly sampled items. In the teaching setting, when given explicit knowledge of the transformation, the strongest model generates demonstrations that teach other models nearly as effectively as human-authored examples. The other models teach considerably less effectively, even when they can apply the provided rules themselves. With an unconstrained interaction budget, most models do not discover substantially better and typically stop after only a few queries. Our results suggest that inferring a rule from given examples and gathering the right examples oneself are distinct capabilities, and that current models may lack the latter even when they have the former.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.