acceptodds
Under review as a conference paper at ICLR 2027

From Relevant Fragments to Sufficient Answers: Verifiable Evidence Chains in Long Dialogues

Abstract

Long-context language models can process large dialogue histories, yet supplying more context does not reliably improve memory-based question answering. We study this through a fixed-panel reader experiment at increasing input lengths and through 50K-token, 100K-token, full-history, and oracle-evidence conditions on three long-dialogue benchmarks. Performance is non-monotonic and dataset-dependent: 50K packages are close to Full on RHELM and LoCoMo Long but substantially worse on EverMemBench, and coverage of Gold Evidence does not guarantee answer correctness. These gaps motivate a hypothesis of evidence sufficiency: when a complete answer is present in a finite textual context, at least one linear chain of source facts, events, and objects joined by verifiable lexical keys supports it. A manual audit of 24 blindly selected questions illustrates chain reconstruction and exposes limitations of official annotations, including dependencies on external metadata.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.