Retrieved but Not Used: Access–Use Dissociations in Long-Context Language Models
Abstract
Long-context evaluations ask whether a language model can find information far back in its input. We ask a narrower question: once a value can be found, can the model use it? We construct partially evaluated programs in which a probe value is defined once and consumed once after a controlled number of intervening lines, with total length and live state held fixed, and we compare three queries at the same position: copy the value (retrieve), add a one-digit literal to it (compute), or add a one-digit local variable to it (combine). Across four Qwen3 models (4B to 32B) at 28K tokens, retrieve stays at 88.5 to 99% after 1,024 lines of separation while compute falls to 37.5 to 76% and combine to 0.5 to 43.5%. On identical programs the use failures sit inside the retrieval successes: Qwen3-8B copies the value on all 200 programs on which it then fails compute (100) or combine (198), and the paired excess of combine over retrieve is 47.5 to 86.5 points on the three models tested. Context structure matters in a scale-dependent way. At a fixed distance, growing the context from 3.7K to 28K tokens shows no detectable load effect on retrieve or compute at 8B, 14B or 32B, while combine falls from 95 to 66% at 32B; the far-distance collapse occurs at constant length. Across families the dissociation is a boundary rather than a universal: Llama-3.1-8B shows the one-operation gap (19 points) but not the two-variable collapse, and Gemma-3-12B shows both (37 and 61 points) where its position-dependent retrieval holds. A 2,048-token generation budget returns compute to 86 to 99% on the same items with or without an instruction to reason, while the question format at short budgets answers nothing; what the budget buys is not identified. A pre-specified decision rule based on a log-odds interaction returned no-go; under ceiling-level retrieval that estimator tests a different scale from the verbal hypothesis, so we make no confirmatory claim from it and label the probability-scale analysis post hoc. All evidence comes from one synthetic task.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.