ConceptARC at Scale: Studying Conceptual Understanding in Reasoning Models
Abstract
Large reasoning models (LRMs) perform well on challenging benchmarks, but it remains unclear whether they form generalizable abstractions across conceptually related reasoning tasks. The Abstraction and Reasoning Corpus (ARC) provides a controlled testbed for studying abstraction, requiring solvers to infer abstract transformation rules over two-dimensional grids from a small set of examples. While ARC evaluates rule abstraction within individual tasks, different tasks may share higher-level structure despite requiring different transformationsfor example, several tasks may involve reasoning about objects contained within other objects. Benchmarks such as ConceptARC capture this structure by grouping ARC-style tasks into recurring spatial and semantic concepts, but existing datasets contain only a small number of manually constructed tasks per concept. We introduce ConceptARC at Scale, a generator for creating thousands of new ARC-style tasks targeting a specific concept. Using this generator, we construct and publicly release ConceptARC-100, a dataset of 1,600 new, manually reviewed tasks, with 100 tasks for each of the 16 concepts defined by ConceptARCten times as large as the original ConceptARC benchmark. We evaluate a range of recent LRMs, as well as models trained from scratch on ConceptARC-100, to characterize which concepts they solve reliably and where they struggle. We further analyze the models' internal representations and find clear semantic structure: tasks cluster according to the transformation operations applied to objects, such as recoloring, removal, or translation, even when these transformations are realized through substantially different rules and scenarios across tasks. Together, our results suggest that models form stable abstractions of transformation operations that generalize across diverse tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.