acceptodds
Under review as a conference paper at ICLR 2027

A minimal probe of task modeling in language models

Abstract

Language models use long contexts and score well on knowledge and reasoning benchmarks, but neither says whether they can carry out work whose next step depends on everything that has happened so far, as a research program or a long-running agent does. We call that ability task modeling. Counting identical items is its minimal case, a total to which every item contributes and that no item records; the question is whether a model's limit on counting is a limit of that ability. Here we introduce counting capacity (CC), the length up to which a model's exact count is its majority behavior, and measure it for 205 models in 296 configurations. Counting fails early: every resolved measurement loses count, at a median of 50 items and at best 627, and across release cohorts the median rose 1.6-fold while the median context window grew about 80-fold. As far as the measurements can tell, the limit is not a simple limit of counting: the failures are abrupt, structured collapses rather than careless slips; the capacity shifts severalfold with the encoding while the ordering of models holds; native reasoning gives no median gain across 77 verified contrasts, with large model-specific gains and losses; and a concurrently executed rule task consumes it in both models tested. Mechanistic analyses of six open-weight models locate the count in an internal state that stays readable in magnitude past the boundary, and an intervention that carries the count forward carries, unchanged, the state of an object history. Counting capacity tracks the METR task horizon beyond general capability on the 24 models shared, scales with a ladder of grid-execution tasks, and co-varies across vendor families with the length at which a state-tracking task fails when reasoning is off, a limit the same models escape when reasoning is on and the state can be written down. These findings are consistent with counting measuring, at a few hundred characters per probe, a distinct, silent dimension of capability that a chain of thought masks whenever the state is nameable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.