acceptodds
Under review as a conference paper at ICLR 2027

InversePerturbBench: Candidate-Aware Evaluation for Single-Cell Perturbation Retrieval

Abstract

Inverse perturbation retrieval identifies the intervention behind a measured transcriptional response by ranking candidate interventions, and its success depends on separating candidates rather than on fitting responses on average. We present **InverseperturbBench**, a benchmark that measures how much intervention identity survives in a candidate library, under matched queries, matched candidate sets, and a single scoring rule. Across 41 external query actions and 1,087 candidates, measured responses recover 87.80% of interventions at rank 10, against 19.51% for an archived GEARS predictor. We call this shortfall the *identity gap*. It is not a generalization failure: on the 28 actions the predictor was fitted on, recovery is 85.7% versus 28.6%, and every predicted top-ten hit falls inside that fitted stratum. It is also not an artefact of target-gene readouts, since a target-excluded comparison retains a 16.98-point gap. Predicted first ranks concentrate on a shared translation-related response rather than on candidate-specific detail, yet removing that concentration alone does not recover accuracy. We therefore develop \cari, which rescores a frozen candidate library against the identification objective without retraining any forward model and improves genetic retrieval from 15.17% to 18.83% and drug retrieval from 24.4% to 33.3%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.