acceptodds
Under review as a conference paper at ICLR 2027

Quantifying the Link Between Localization and Context Interventions in Long-Context QA

Abstract

Search-augmented question answering, knowledge-base question answering, and multi-document analysis require language models to identify the right evidence among many documents and use it to produce an answer. Whether models can do this reliably determines whether longer context windows translate into stronger problem-solving capabilities. Existing methods use localization or relevance sig- nals to prioritize context and improve its utilization through attention calibration, document reordering, compression, or pruning. Although prior work reports both localization quality and downstream QA performance, it does not explicitly quan- tify how localization successes and failures jointly determine intervention out- comes. We present a systematic analysis linking localization accuracy to down- stream QA outcomes across multiple matched context interventions. In the actual locate-then-intervene workflow, we group records by whether the localizer se- lects the answer document, measuring the average gain after a hit and the average loss after a miss. A separate gold-document control applies the same interven- tion directly to the answer document to measure the intervention’s improvement potential. We further show that document-localization signals concentrate in a small number of attention heads, and select these heads using questions disjoint from evaluation to connect localization quality to intervention outcomes. Experi- ments across multiple models and datasets show that localization and intervention based on a small number of attention heads improve answer-hit accuracy by 4.46 percentage points on average and yield positive effects in the large majority of configurations. Gold-document controls further show substantial remaining im- provement potential across the tested interventions. On unseen models, using only a small labeled set, our method predicts whether an intervention will help or hurt with 78.1% accuracy, compared with 58.8% for the better constant policy, without using held-out full-QA outcomes at prediction time. These results turn localiza- tion accuracy from an isolated intermediate metric into a basis for explaining and predicting the effectiveness of context interventions, providing a unified view of how such methods can be analyzed and transferred.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.