Position Bias Is Not Monotone: Needle Recovery Profiles Over Real Long Documents
Abstract
Long-context large language models are routinely evaluated using synthetic "needle-in-a-haystack" retrieval tasks, which assume a monotonic drop-off in retrieval success as context length increases or as positions shift from boundaries to the interior. We demonstrate that this assumption fails on real long-document distributions (such as multi-chapter Gutenberg texts). By evaluating retrieval and recovery profiles under controlled position manipulations, we find that retrieval accuracy exhibits a distinct U-shaped or non-monotone profile rather than simple exponential decay. Interior retrieval performance does not uniformly degrade as a function of relative depth; instead, token-level salience, lexical interference, and local attention crowding dictate recovery more aggressively than sheer distance from context windows. We provide a rigorous diagnostic framework for evaluating long-context architectures on natural corpora rather than synthetic filler, showing how standard benchmarks over-penalize models due to artificial prompt artifacts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.