acceptodds
Under review as a conference paper at ICLR 2027

A benchmark score is not a model property: failure composition changes across operating regimes

Abstract

Language models are typically evaluated through aggregate metrics such as accuracy, which count how often a model fails without recording how it fails. A rising benchmark score can therefore conceal whether an intervention actually improves reliability or merely changes the type of errors a model makes. We study this question with a framework that decomposes failures into various mutually exclusive mechanisms, such as a wrong action (a decision error) or a misreported resulting state (a state error), and tracks how their composition changes across operating regimes and interventions. Using an exact-simulator benchmark for multi-step reasoning, with no model-based judging, we analyze failure composition across 13 models from 9 families, including open-weight and proprietary systems. We further evaluate five interventions across reasoning, decoding, and representation changes, measuring the aggregate performance and failures movement. Our analysis reveals two consistent patterns. First, failure composition changes systematically across models and capability regimes. Models with similar aggregate scores can fail through different mechanisms, and across eleven models the dominant failure type shifts with the model's reliable operating horizon (Spearman ρ = 0.79). Second, an intervention can eliminate its targeted failure while other failure modes absorb the errors, sometimes without improving overall reliability. In Qwen3-4B, position-based representations reduce state errors from 0.506 to 0.006 while introducing a new formatting failure, and self-consistency turns half of Qwen3-1.7B's would-be decision errors into state errors. These findings show that aggregate metrics alone provide an incomplete picture of reliability. We argue that evaluations should measure how failures change under scaling and intervention, as well as how many remain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.