Gather, Don't Scatter: Trainable Placement Robustness for Multi-Hop Reasoning in Multi-Document Long-Context LLMs
Abstract
Language models are increasingly given multiple documents in context to answer questions that require external, specialized, or up-to-date information. However, their answers can depend on how those documents are arranged. We examine this dependence in multi-hop question answering, where every document, gold or distractor, is placed directly in the context and each gold document supplies one step of the reasoning. We separately examine three aspects of document ar- rangement: the order of gold documents in the reasoning chain, their position in the context, and whether they are grouped together or scattered among distractors. Across two 8B base models, whether the gold documents are together or scattered matters most. Keeping them together helps consistently wherever the block sits, while chain order and exact block position have smaller effects. We then vary only the arrangement of the fine-tuning data. Training on examples with scattered gold nearly eliminates placement gaps and improves pooled accuracy by 20–24 percentage points, whereas training on a fixed gold-at-the-front layout induces se- vere brittleness when the gold moves. These findings hold across three datasets, with contexts up to 32k tokens, under harder retrieved distractors, and in long- form generation with citations. The main result also holds for a 27B model from a third family. Diagnostic probes further show that many remaining failures occur even when the individual supporting facts are accessible, identifying multi-hop composition as an important residual bottleneck.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.