PROVER: A Claim-Centric Framework for Auditing Reasoning Traces in Cellular Perturbation Analysis
Abstract
In the context of AI for Virtual Cells (AIVC), cellular perturbation models increasingly generate natural-language analyses alongside predictions of cellular responses. While these analyses promise biologically grounded interpretability, their validity is rarely examined beyond final prediction accuracy. We introduce (erturbation easoning utput ifier), a claim-centric hierarchical framework for auditing reasoning traces in cellular perturbation analysis. PROVER reconstructs each trace as a structured argument graph centered on Claims and evaluates it along two dimensions: and . Cogency assesses whether premises are supportable, evidence is relevant to target Claims, and inferential connections provide sufficient support for their conclusions, while Completeness measures coverage of task-relevant analytical perspectives. We apply PROVER to a 600-case benchmark spanning eight perturbation-analysis systems, examine three LLM audit backbones and holistic-judge baselines, and complement these analyses with expert feedback. The audits flag unsupported assumptions, context-misaligned evidence, and insufficiently supported inferences that answer-level metrics do not assess. Our results show that evaluating biological reasoning requires assessing not only prediction outcomes but also the evidence and reasoning used to justify them. We release the framework, auditing protocol, benchmark metadata, and audit artifacts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.