acceptodds
Under review as a conference paper at ICLR 2027

Ace Olympiad Math, Fail Simple Arithmetic: Systematic Capability Unreliability in LLMs

Abstract

Frontier language models now reach gold-medal-level performance at the International Mathematical Olympiad, yet still make systematic errors on simple tasks like arithmetic. Methods that generate and search over new inputs can expose such failures, but patterns discovered on adaptively selected failures may not generalize, and predictive patterns need not identify what causes them. We introduce a framework that separates failure discovery from inference. Starting from programmatically verifiable task generators, we search for failures, grow local failure clusters, and interpret each cluster with an executable hypothesis. We accept a hypothesis only if it predicts failures on fresh inputs under input-level, false-discovery-rate-controlled validation; designed contrasts then test whether individual input properties actually change the failure rate. Across six capabilities and 17 language models, including five frontier models, failures are common and strongly clustered, yet adaptive search provides little advantage over uniform sampling at the failure rates we study. More importantly, only 19 of 107 search-derived hypotheses survive validation. In simulations with known ground truth, drawing fresh inputs eliminates 141 of 144 false acceptances near the validation threshold. Designed contrasts further reveal failure-inducing structure that observational validation alone can miss: carry cascades cause addition failures in open-weight models and in three of five frontier models. These results turn failure search from a procedure for collecting errors into a method for making statistically tested claims about where language-model capabilities are unreliable, and suggest that LLM evaluation should move beyond aggregate scores and failure collection toward such validated claims.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.