acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Cell and Drug Representations for Generalizable Drug Response Predictions

Abstract

Foundation models provide increasingly sophisticated representations of cells and drugs, with promising applications in precision oncology and drug discovery. Yet it remains unclear when these representations improve drug-response prediction or generalize across drugs, cellular contexts, and datasets. We systematically assess the effects of cell and drug representations, predictive architectures, and dataset characteristics by benchmarking five cell representations, six drug representations, and three predictive models across two multimodal single-cell settings: a large, controlled shared-label dataset with uniformly generated transcriptomes and standardized cell line–drug response labels, and a smaller, heterogeneous cell-specific-label setting with finer-grained response annotations. We evaluate increasingly stringent regimes, from random splits to zero-shot prediction of unseen cell line–drug pairs, drugs, and cellular contexts. Across both settings, performance depends more strongly on the cell representation and predictive model than on the drug representation, while pretrained encoders do not consistently outperform simpler alternatives. Under cell line–drug pair zero-shot evaluation, the best MCC reaches 0.36 in the shared-label setting and 0.17 in the cell-specific-label setting. Generalization degrades sharply with stringent splits compared with random splits: in the shared-label setting, the average within-drug MCC decreases from about 0.89 with a random split to 0.18 with cell line–drug pair zero-shot. In the cell-specific-label setting, it decreases from 0.62 to 0.03. Conclusions also depend strongly on the evaluation axis; for example, shared-label cell-line zero-shot yields MCC 0 when evaluated across cellular contexts but 0.52 when aggregated across drugs. Multimodal models likewise do not consistently outperform single-modality baselines, indicating that much of the predictive signal can often be captured by one modality alone. These results highlight the importance of realistic zero-shot evaluation, complementary aggregation views, and strong simple and single-modality baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.