Research Agents Get Lost in Evolving Evidence
Abstract
A research question may ask what was supported by the evidence available at a particular date. Scholarly corrections and revisions provide documented cases in which that answer changes. We study this requirement as temporal grounding and introduce Time-conditioned Scholarly Conflict Question Answering (TSC-QA), a benchmark of 200 question pairs across eight scientific fields. The questions in each pair differ only in their evidence cutoff and require different answers. Strict-pair accuracy requires both answers to be correct. To test this requirement across different ways of conducting research, we evaluate fixed answerers, planned workflows, and systems that control their own research loop. These include Qwen, Llama, Gemma, DeepSeek, and GPT models, as well as the complete Codex-Luna, Codex-Terra, and Codex-Sol systems. When systems must find or recall the evidence themselves, the highest pair accuracy is 35.50%, and the delegated research systems perform better at the later cutoff. Providing claims organized by source and public date improves strict-pair accuracy more than providing claim summaries alone. Models can recognize a supplied claim and answer the later question correctly while still missing the earlier answer. Response analysis connects these results to claim interpretation, source dating, and evidence use. TSC-QA tests whether research agents can recover the answer supported at each point in a changing evidence history.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.