acceptodds
Under review as a conference paper at ICLR 2027

AXIOMSHIFT: Altering Axioms to Evaluate LLM Reasoning At Knowledge Boundaries

Abstract

Mathematical reasoning has become one of the central proving grounds for state-of-the-art AI, with recent systems solving several long-standing open problems. However, it remains unclear whether a model’s answer follows from the given premises or prior knowledge encoded in its parameters, or if the model even recognizes when the provided premise lacks enough information to determine an answer. On conventional benchmarks, these differences are indistinguishable. We introduce AxiomShift, a 700-question benchmark that isolates two distinct boundaries of model knowledge. First, it distinguishes supplied from parametric knowledge by perturbing a single axiom in each evaluation world, forming counterfactuals that test whether a model answer is attributable to prior knowledge or the given premise. The second boundary tests whether a model recognises and respects what can be known from given premises, through withholding or contradicting information required for the derivation, which produces genuinely unanswerable questions. Across 26 models, we identify four distinct model failure types. As model accuracy increases, simple numerical errors decline. However, prior leakage itself demonstrates a sharp split, with leakage on answerable problems decreasing (12.6% to 1.9%), while rising on unanswerable problems (12.1% to 15.5%). Taken together, our findings suggest model scaling improves ability to follow supplied premises, however, does not reliably improve recognition of knowledge boundaries.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.