acceptodds
Under review as a conference paper at ICLR 2027

EvidenceWeave: Compiling Time-Bounded Evidence into Inspectable Mechanisms and Discriminative Tests

Abstract

One omitted contradiction can reverse the rescue experiment chosen from an otherwise well-formed mechanism graph. EvidenceWeave keeps that failure visible by compiling the path from each causal edge to its source span, accepted or uncertain state, unresolved prerequisite, rival mechanism, and outcome-separated test. The accompanying benchmark contains 312 CRISPR perturbation cases and 4,210 expert-identified support or contradiction opportunities, including 715 contradictions, from literature published before January 2022. Chronology is enforced before retrieval ranking: fixed-volume future-document injection yields 0.089 leakage and 8.3% support inflation at 20%, whereas the date-filtered pipeline has zero detected leakage. Against a typed Llama-3.1-70B control matched on training objects, retrieval, schema, and decoding, EvidenceWeave lowers unsupported-edge rate from 0.246 to 0.142 while raising contradiction coverage from 0.302 to 0.478 and evidence-opportunity recall from 0.601 to 0.672; 299 additional annotated opportunities reach reviewers with similar accepted-edge mass. A separately blinded 72-case annotation panel preserves the edge ordering, while a 48-case human target yields written-outcome-separation AUC of 0.706 versus 0.584 for the control. Counterbalanced review also falls from 11.1 to 8.7 minutes, with faster review in 49 of 60 pairs. The resulting model-selection criterion rewards evidence exposure and adverse-evidence recovery only when asserted edges remain supportable and the linked object reduces expert reconstruction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.