acceptodds
Under review as a conference paper at ICLR 2027

Discover, Falsify, Revise: Auditing Input-Use Claims from Source Code to Predictive Contribution in Agent-Discovered Cell Models

Abstract

AI virtual cells aim to predict cellular responses to specified interventions, yet held-out predictive performance alone does not establish use of the supplied perturbation information. This prediction–claim gap is especially consequential in agentic model discovery, where language-model agents generate and revise predictors using predominantly score-based feedback. We introduce CellAudit, which audits registered input-use claims through three distinct questions: whether the input can enter the cited computation, whether fitted predictions depend on it, and whether that dependence improves prediction of observed responses. On a paired morphology–transcriptomics perturbation benchmark (BBBC047), an agent-selected predictor achieves a mean held-out Global Pearson correlation coefficient (PCC) of 0.3153 yet remains exactly invariant to compound replacement; a control-profile-only predictor reaches 0.3142. Source inspection identifies a compound-query pathway blocked by singleton key–value attention; the invariance persists after refitting with physically disjoint control wells for inputs and target references. In a stratified audit of 48 generated candidates across two linked tasks, 47 change predictions under compound replacement on both held-out folds, but only 20 show target-loss gains with intervals above zero on both folds. On BBBC047, falsification-guided revisions recover positive mean compound contributions while retaining predictive gains over the control-profile-only baseline. In matched sci-Plex searches, audit-enriched feedback shows higher mean held-out predictive performance and larger mean compound and dose contributions across five paired trajectories, although paired intervals span zero. Refitting fixed designs on an independently acquired cohort further shows predictive generalization need not imply generalization of input-use claims: dose contribution persists, whereas support for compound identity does not persist. CellAudit therefore adds a falsification layer to agentic model discovery, moving from generate–score–revise toward discover–falsify–revise.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.