acceptodds
Under review as a conference paper at ICLR 2027

Does Recursive Knowledge Construction Help Agents Verify Consensus Claims in Science?

Abstract

Recent scientific agents build knowledge recursively by retrieving evidence, constructing an intermediate state, and judging whether that state contains sufficient information. These agents are evaluated end-to-end, usually without ground truth, so the effect of each design choice is unknown. We examine two design axes: (1) what the recursion builds (the state of evidence) and (2) what controls it (the judge that decides when to stop). To measure them, we define consensus scientific claim verification and measure these effects on this task, where a system must establish a claim's consensus stance from the open literature. We build a new consensus claim dataset, ProClaim-eval, including 718 claims with SUPPORT and REFUTE gold labels from five expert-curated databases in molecular biology, nuclear physics, and astronomy. Our instrument, ProClaim, persists an evidence state outside the model's context, so each axis can be studied with the other held fixed. Against verification systems and general-purpose agent harnesses, ProClaim ranks among the top three, with a significant advantage over the harnesses. We find that the evidence state improves performance over in-context accumulation, and it also matters what is stored in it. In addition, a trained judge that reads the evidence state performs better than fixed-round stopping. A judge trained separately and never exposed to ProClaim-eval achieves the highest agreement with the gold labels on three of five subsets, where it matches the off-the-shelf Claude SonnetĀ 4.6 judge and outperforms Claude HaikuĀ 4.5. Our results on ProClaim-eval support recursively maintaining a structured evidence state outside the model's context as an effective design for verifying claims in scientific agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.