The Context Counts: Characterizing LLM Counting Failures Across Input Regimes
Abstract
Counting is an essential skill underpinning complex forms of quantitative reasoning, yet large language models (LLMs) remain unreliable even on seemingly simple counting tasks. Prior work has shown that contextual cues and training frequency biases affect counting performance. Building on this literature, we characterize not only how often LLMs miscount, but also the magnitude and direction of their errors across two regimes: non-contextual counting (tallying items in line-separated lists) and contextual counting (a filter-then-count operation over unstructured text passages). Our experiments reveal three findings: (1) miscount rates rise with target count, although most errors are relatively small and undercounting predominates; (2) greater list diversity improves performance, even when the added variation comes from case perturbations on a fraction of instances, interspersed sporadically in an otherwise homogeneous list; and (3) similar miscount rates across degrees of context corruption conceal distinct error patterns: GPT 5.2 makes larger relative errors on partially corrupted than fully corrupted contexts, while Claude Opus 4.5 begins to refuse to perform the counting task on contexts with corruption rates as low as 25%. Together, these results show how changes to input structure affect counting errors in ways that miscount rates alone do not capture. We hope this behavioral account provides a basis for future mechanistic studies of LLM counting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.