acceptodds
Under review as a conference paper at ICLR 2027

Beyond Context Length: Characterising the Difficulty of Long-Context Reasoning

Abstract

Long-context language models support ever longer context windows, and evaluations typically treat context length as the main indicator of difficulty. However, context length alone cannot explain why models fail on long-context tasks that require reasoning over complex information. Contexts of equal length can pose very different challenges: a task may be solvable from a single local passage, or it may require integrating evidence distributed across the full context. To characterise this difficulty systematically, we propose quantitative, sample-level definitions of two evidence-related properties, *scope* and *dispersion*, and further extend this characterisation with *reasoning depth*. To study how these three properties contribute to model failures, we construct *LURCH*, a multi-domain question-answering benchmark in which each question is provided with its answer and supporting evidence, enabling all three properties to be quantified. Controlled experiments on both proprietary and open-weight models show that: (1) at fixed context length, all our three properties are strongly negatively correlated with model performance; (2) as context length increases, performance degrades sharply on samples with dispersed evidence, whereas it remains relatively stable on samples with localised evidence; and (3) difficulty compounds across properties, as questions with moderately high values on all three are harder than those with an extreme value on just one. These findings establish evidence scope, evidence dispersion and reasoning depth as measures of long-context reasoning difficulty that complement context length.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.