FalsifyVC: Error-Controlled Reliability Testing of Virtual Cell Models across Comparator Authorities
Abstract
Virtual cell predictors are judged by held-out accuracy on unseen perturbations. Their added value depends on the comparator's selection information and rule, which existing reliability tests take as given. FalsifyVC indexes verdicts by this comparator authority, separating the reliability threshold, margin and run budget. Exact binomial tests count training runs whose excess over the comparator clears a registered margin. On two Replogle screens, no GEARS run clears the margin against any comparator across twelve fresh target splits with its reference preprocessing, or against fixed pseudobulk and a held-out per-target envelope on the registered split. STATE and CPA are refuted outside their supported regime, STATE also with training-data gene features. A training-only selector yields positive GEARS mean excess on the registered split but is weaker than training-selected ridge, and the excess vanishes on fresh splits. On an independent screen GEARS clears the margin in none of thirty runs, on one split or across thirty. A registered positive control clears the margin in all thirty runs and is supported, then refuted once an additive baseline is registered. The results refute the specified added-value claims under these conditions and show that the rule can support a large advantage over registered comparators.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.