acceptodds
Under review as a conference paper at ICLR 2027

SenseMath: Do LLMs Have Number Sense? Evaluating Shortcut Use, Judgment, and Generation

Abstract

Large language models often default to step-by-step computation even when efficient numerical shortcuts are available. This raises a basic question: do they exhibit number sense in a human-like behavioral sense, i.e., the ability to recognize numerical structure, apply shortcuts when appropriate, and avoid them when they are not? We introduce SenseMath, a controlled benchmark for evaluating structure-sensitive numerical reasoning in LLMs. SenseMath contains 4,800 items spanning eight shortcut categories and four digit scales, with matched strong-shortcut, weak-shortcut, and control variants. Guided by Cognitive Load Theory, we manipulate intrinsic load through digit scale and shortcut availability while controlling extraneous load through a consistent item type and format. Following Bloom's Taxonomy, SenseMath supports three evaluation settings of increasing cognitive complexity: Shortcut Use, Applicability Judgment, and Problem Generation. Our evaluation across six model configurations reveals a consistent gap between executing shortcuts and understanding their applicability. Explicit number-sense prompting yields accuracy gains of up to 23.0 percentage points on shortcut-amenable items, yet under standard chain-of-thought prompting, shortcut usage on four-digit strong-shortcut items ranges from 26.3% to 58.8%. Models systematically overgeneralize shortcut applicability, with judgment accuracy ranging from 52.5% to 55.9%, and struggle to generate valid strong/control problem pairs, with only 0–21.9% passing all six automatic checks. Together, these results suggest that current LLMs exhibit procedural shortcut fluency without consistently demonstrating the structural understanding of when and why shortcuts work that underlies human number sense.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.