acceptodds
Under review as a conference paper at ICLR 2027

How Many Measured Genes Is a Prior Worth? Unseen-Perturbation Prediction as Decoding

Abstract

Models of transcriptional responses to genetic and chemical perturbations rarely beat the mean training response on unseen perturbations. We ask what a small measurement of the new perturbation would add, and whether its value can be predicted. Treating prediction as decoding, we note that without perturbation-specific information beyond the training set no architecture beats the mean response, predict the expected held-out error of linear measurements (a few genes or cells, or another cell line's response) from training perturbations alone with the classical risk of least squares, and bound combinations of training responses by a coverage floor. On five screens from three studies, the best tested zero-shot predictor recovers 7% to 30% of the recoverable signal and larger tested decoders at most 1.0 percentage point more, whereas four greedily chosen genes of the new perturbation (a simulated targeted readout) recover 33% to 75%; the best prior is worth about half of the first greedy gene on four screens (0.4 to 0.6 gene equivalents) and 3.4 genes on the fifth. With eight or more training perturbations per measured dimension, predictions track held-out errors (slope 0.97, , optimism 0.03) and choose panel sizes better than naive estimates, but are optimistic by about 0.11 below 100 training perturbations. Withholding a pathway raises the coverage floor from at most 13% to at least 33% of the held-out signal, and genes, priors and predictions fail; covering a quarter of it brings genes and predictions most of the way back.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.