Beyond Lost in the Middle: Causal Locality in Tool-Using Agent Contexts
Abstract
Long-context evaluations usually move a relevant document while treating the remaining input as exchangeable. Agent histories are different: a tool call and its result form an atomic transaction, and later transactions can depend on values or decisions produced earlier. Reordering an agent context may therefore gain favorable absolute position while destroying causal locality. We introduce a paired intervention testbed that varies the producer position and the amount of unrelated evidence inserted inside a critical dependency chain, while holding the evidence multiset fixed and preserving a valid topological order. Across 24,144 greedy-decoded calls to two open-weight model artifacts, the separation contrast is -4.40 percentage points for GPT-OSS 20B, whereas Gemma 3 27B has no changed binary outcomes in the paired cells. Their equal-weight arithmetic mean is -2.20 points (95% trajectory-bootstrap CI: [-4.05, -0.23]), but is not a cross-model population estimate. Because the original separation treatment also moves the consumer endpoint, this pilot does not identify a position-independent locality coefficient. We find no reliable front-versus-middle benefit or pre-registered interaction. An exploratory end-pinning heuristic improves GPT-OSS by 4.63 points over chronological order, but has no effect on Gemma. These results expose capability, endpoint-position, and run-reliability checks that are necessary before asserting a universal agentic “lost-in-the-middle” rule or deploying an optimized placement policy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.