InvariantBench: Can Large Language Models Exhibit Inherent Reasoning Consistently Across Equivalent Transformations?
Abstract
Reasoning is often attributed to large language models (LLMs), yet it remains unclear whether they operate over underlying semantics or rely on surface-form patterns. Existing benchmarks evaluate correctness on fixed problem instances, but overlook a fundamental property of reasoning: invariance under semantics-preserving transformations. If a model truly understands a problem, its predictions should remain consistent across equivalent representations. We introduce ***InvariantBench***, a benchmark of 1159 seed problems invariant forms, spanning 16 tasks across 3 reasoning families and 12 fine-grained invariance axes. Each problem is paired with multiple semantically equivalent variants under a strict invariance contract, enabling evaluation beyond accuracy to measure consistency across representations. Experiments on 15 frontier and open-weight LLMs reveal a persistent invariance gap: invariant reformulations reduce accuracy by up to 30% for strong models, while the gap between solving at least one versus all four variants reaches 60%. Full consistency remains below 5% for most open-weight models, compared with 48% for the two strongest models and 98.6% for human experts. These results show that high base accuracy substantially overestimates reasoning ability, and establish invariance as a necessary axis for evaluating and improving robust language understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.