acceptodds
Under review as a conference paper at ICLR 2027

Eclipse: Joint Risk Control for LLM Output Recovery and Evidence Completeness

Abstract

Selecting fewer retrieved passages can reduce the amount of text needed to inspect a language model's existing output. However, reproducing that output from fewer passages does not establish that the supporting evidence remains available. We introduce Eclipse, a framework for evaluating output recovery together with evidence completeness, defined as retaining all benchmark-annotated supporting passages. Its search keeps a small initial set of retrieved passages and adds passages chosen by the likelihood of the original output. It stops when the model reproduces that output or reaches a preset size limit. We then evaluate the fixed search procedure on separate calibration samples. These provide upper confidence bounds on recovery failure across queries and missing evidence among returned sets. Certification requires both bounds to meet preset limits under the stated independence and sampling assumptions. On held-out 2Wiki queries, search with Qwen and a supervised retriever satisfies both requirements for 94.1% of queries, exceeding retrieval-order addition by 9.2 percentage points. This configuration also meets the preset limits for both calibration bounds. Across our experiments, search mainly helps when the initial passages already contain the supporting evidence; recovering the output often leaves missing evidence unresolved. Jointly checking output recovery and evidence completeness therefore helps assess smaller passage sets, while neither establishes answer correctness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.