acceptodds
Under review as a conference paper at ICLR 2027

StructScope: Can LLMs Maintain Structural Scope in Long Contexts?

Abstract

The rapid growth of structured information across long contexts, such as scientific documents, software artifacts, and harness-managed contexts creates a need for models that can determine not only which information is relevant but also which structural scope makes it admissible. However, existing long-context benchmarks primarily evaluate models' capability to process text information with explicit structured information as fixed background. To bridge this gap, we introduce , a novel benchmark designed to evaluate models' capability in , i.e., applying query-specified scope constraints when selecting and attributing information. Although models show strong structural reasoning when evaluated on compact structural parse trees, our experiments reveal that their performance degrades sharply as the same structures are embedded within longer contexts. Further fine-grained experiments identify scope-constrained selection, particularly within nested structures, as the key bottleneck. Our findings highlight structural scope maintenance as an important dimension of long-context capability, motivating dedicated evaluation, training, and scope-aware context management in harness systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.