Found but Not Written: Understanding and Reducing Evidence Attrition in Deep Research
Abstract
Deep Research agents can retrieve relevant evidence yet omit details readers need. We study this evidence-to-answer attrition on 132 bilingual DeepResearch-Bench II (DRB2) tasks, where a selector flags candidate omissions in about 80% of reports. In controlled completion experiments, targeted additions fully express more omitted findings than untargeted additions, without improving mean report quality. We introduce Evidence-Grounded Residual Completion (EGRC), which organizes the answer around the request and revisits every original omission to retain content already expressed, revise a passage, or justify deferral. On 36 DRB2 test tasks, requiring a decision for every omission adds 3.9 quality points over freely selecting from the same omission list, with the same organized starting answer and nearly equal final length. The full method improves quality by 4.7 and 5.0 points over untargeted completion and two-pass self-checking, respectively, while recovering more omissions. On 23 scholarly tasks constructed following DeepScholar-Bench, EGRC improves recovery by 33 percentage points and reference-content coverage by 8.8 points over matched autonomous revision at nearly equal length. The content advantage persists under report-only assessment, and blinded readers favor EGRC over completion baselines. Explicit finding-level decisions improve recovery and report quality beyond answer organization and access to an omission list.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.