acceptodds
Under review as a conference paper at ICLR 2027

Risk-Controlled Biomedical Claim Verification over Heterogeneous Evidence Sources

Abstract

Verifying a biomedical claim is adjudication across an ecosystem rather than lookup in a corpus. One assertion that a compound inhibits a target is recorded at once in abstracts, curated pharmacology records, interaction graphs and solved structures, each with its own provenance and update cadence, and these records disagree for structural rather than accidental reasons: a curated entry may assert an interaction that a later abstract negates under a different assay. Existing resources assume one corpus with one gold rationale, so they cannot express such conflict, and the verifiers built on them return unauditable labels alongside confidences that do not bound error. We introduce BioMSVerify, claims over four types with gold evidence spanning at least two source modalities, with conflicting sources, and Evidict, a verifier fusing a generative energy over typed evidence, a reliability-weighted and direction-preserving contradiction score, and a bidirectional process reward model multiplying each step's grounding by its value-to-go, then turning the fused score into a commit-or-abstain decision by group-conditional split conformal prediction. Three measurements motivate the design. Closed-book PubMedBERT reaches macro- overall and on multi-hop drug–target–pathway–phenotype claims, so accuracy cannot certify grounding. Given an identical evidence pool, a frontier reasoning LLM reaches while a M retrieval-grounded encoder reaches , so scale is not the bottleneck. And confidence is not risk: committing on temperature-scaled scores leaves a realized error of on conflicting-source claims, twice a target, whereas the conformal procedure keeps realized risk at or below for every claim type and on the conflict subset at three target levels. Evidict improves macro- by points on conflicting-source claims over the strongest baseline, is the most robust tested system to five counterfactual perturbation families, and costs as much at inference as GPT-4o.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.