Hitting the Right Targets: Simple Pseudobulk Predictors for Chemical Perturbation Prediction
Abstract
Predicting transcriptional responses to chemical perturbations can accelerate drug discovery. Single-cell perturbation models are increasingly complex, yet they are typically evaluated on aggregated pseudobulk profiles, where simple baselines remain competitive. We revisit chemical perturbation prediction at pseudobulk resolution with a lightweight MLP on frozen compound and cell-line embeddings, and benchmark twelve compound representations spanning molecular structure, morphology-aligned phenotype, and drug targets. On sci-Plex3 and Tahoe-100M, the predictor outperforms learning-free baselines and four single-cell models (ChemCPA, BioLord, Doloris, and PerturbDiff) on most metrics while using up to 55 fewer parameters. To understand why, we examine the two ingredients of the predictor: its prediction target and its compound representation. Aggregation yields a less noisy target, and retraining a representative single-cell model (ChemCPA) on pseudobulk profiles narrows, but does not close its gap to our model. On the representation side, more sophisticated molecular embeddings bring little consistent gain; instead, performance tracks how closely the geometry of a representation aligns with that of the transcriptional responses. These results suggest that predictive performance depends not only on model complexity but also on the data, and position pseudobulk predictors as strong baselines and potential priors for single-cell models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.