The Construct Is Part of the Target: Rethinking Perturbation-Model Evaluation
Abstract
Reagent-level variation is biologically established, but perturbation-model benchmarks typically evaluate predictions at the gene level. We ask whether model evaluation changes depending on which reagent is used to perturb a gene and whether that reagent is available in source data. All 2,053 genes observed in all four core CRISPRi panels reuse at least one complete construct across all four contexts, so cellular context holdout alone does not imply construct holdout. Alternative constructs produce recurring response differences beyond matched cell resampling; external CRISPR KO and CRISPRi libraries also show differences beyond same-guide resampling. Holding predictions fixed, changing the realized target changes absolute model scores; released State predictions also show sensitivity to target realization beyond same-construct resampling. Under global construct exclusion and matched cell counts, all 16 comparisons show higher mean Pearson when the target construct is present in source data than when it is withheld. Construct transfer also changes the relative MAE performance of the two fitted predictors, while same-construct resampling does not reproduce the shift. Their mean Pearson ordering also reverses descriptively in all four contexts. A State architecture trained from scratch under the same exclusion design shows material Pearson degradation in RPE1 and Jurkat. Perturbation-model evaluation therefore requires specifying both target realization and which constructs are present in source data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.