acceptodds
Under review as a conference paper at ICLR 2027

Beyond Evidence Availability: Interference and Reasoning Diagnostics in Long-Context QA

Abstract

Longer contexts let models access more information, including potentially irrelevant or misleading text. We examine how background content affects question answering across context lengths and reasoning settings. Six models at two reasoning settings answer questions requiring information from several documents, with all annotated supporting passages retained in inputs of about 25,000 or 112,000 tokens. Stage 1 compares background sampled from other questions, related passages filtered to exclude answer mentions, and similar related background with a short passage supporting a designated wrong answer. The passage averages less than 0.2% of the input. We first measure the accuracy loss when sampled background is replaced by related text. We then measure the additional loss under similar background containing the misleading passage. This second loss is larger in nine of 12 model–reasoning combinations, with 95% confidence intervals for the difference above zero. Stage 2 examines cases that the same model answers incorrectly with little or no reasoning but correctly with higher reasoning. Keeping the full input, we provide the lower setting with document numbers from the successful higher-setting answer, numbers of the annotated supporting documents, or the first set plus guidance on using the evidence. On selected cases passing guidance checks with valid responses to all three prompts, the guidance-containing prompt exceeds the prompt using the same document numbers alone by 25.1–39.3 percentage points in five of six models, with positive 95% confidence intervals.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.