Why Do Sophisticated Perturbation Prediction Models Not Outperform Learning-Free Baselines?
Abstract
Why are modern perturbation-prediction models often not much better than simple learning-free baselines? This phenomenon has been widely observed in genetic perturbation benchmarks. However, its interpretation in chemical perturbation prediction remains unclear. Using sci-Plex3 as a primary case study, we show that simple pseudo-bulk-level baselines, together with two new single-cell-level learning-free baselines, remain highly competitive, even against modern diffusion-based models. We then propose a three-axis diagnostic framework comprising Input Use, Response Scaling, and Stratified Performance, which evaluates whether perturbation information is used, whether predicted responses scale with the strength of the observed cellular responses, and where along the perturbation effect distribution models succeed or fail. Applying the framework to ChemCPA, CondOT, Doloris, and PerturbDiff reveals that failure to outperform learning-free baselines can arise from two distinct causes: limited perturbation signal in the data or limited utilization of that signal by the model. Independent evaluation on Tahoe-100M further confirms that the competitiveness of learning-free baselines is dataset- and metric-dependent: ChemCPA remains below the perturbation reference, while PerturbDiff exceeds it on all pseudo-bulk metrics but not consistently on distributional metrics. Our results suggest that comparing performance against learning-free baselines alone is not a reliable indicator of model quality. Instead, our diagnostic analyses provide mechanistic insights into why models succeed or fail, help distinguish data limitations from model limitations, and guide the development of more robust perturbation-prediction models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.