K-PRS: Zero-Training Kernel Regression with Frozen Foundation-Model Embeddings for Cold-Start Drug Response
Abstract
Machine learning models for drug response prediction frequently fail to generalize to novel compounds under cold-start evaluation. We show that this failure stems from shortcut memorization: conventional architectures memorize historical compound identities during training, causing performance on unscreened molecules to collapse below naive cell-line baseline averages. To eliminate this vulnerability, we introduce a closed-form kernel estimator over frozen biological and chemical representations, gated by a positive-semidefinite prior that restricts evidence propagation strictly to compounds sharing identical mechanisms of action. By avoiding parameter optimization on response labels, the model is structurally prevented from memorizing entity identities. Across two independent cancer screening benchmarks, a matched five-fold protocol, and synthetic tasks with controlled identity availability, our non-parametric estimator consistently outperforms parameterized baselines. Furthermore, because predictions are formulated as closed-form weighted averages, the estimation weights inherently provide an exact evidence trail, calibrated uncertainty intervals, and verifiable robustness certificates from a single unified computation without auxiliary surrogates. All algebraic claims are formally machine-checked. Our findings establish both a practical predictor and a broader inductive principle: when entity identity is absent at deployment, explicitly refusing to learn it yields substantially stronger out-of-distribution generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.