acceptodds
Under review as a conference paper at ICLR 2027

Making Every Experiment Count: Analogical Molecular Property Prediction

Abstract

Evidence for therapeutically relevant molecular properties is diverse, spanning cell-based experiments, animal studies, and human trials. How can machine learning models utilize such indirect evidence—measurements on related molecules, often from different assay contexts—when predicting a property of a molecule that has never been tested? We introduce an analogical-reasoning approach that, rather than predicting solely from structure, aggregates such evidence from the literature and reasons about its relevance. The core component is a property-transfer model, trained on literature evidence, that scores how likely a measurement on one molecule is to carry over to another. A context optimization procedure then selects evidence maximizing diversity, relevance, and transfer likelihood, which frontier LLMs can use to predict by analogy. We assemble a dataset of direct and indirect safety and ADME evidence mined from a 25M-paper corpus and use the direct measurements to build literature-derived benchmarks for six tasks. With Gemini 3.8 Flash, our system reaches a mean macro-F1 of 0.736 on the literature set and 0.815 on the de facto TDC benchmark, exceeding the best ML baseline on each benchmark—including MiniMol trained multi-task on the same indirect evidence—by 11.9 and 11.2 percentage points, respectively. Against naive retrieval with the same backbone—the comparison that isolates our evidence selection—the gains are 2.4 and 4.5 points with Gemini, and 3.4 and 8.8 points averaged across the three backbones, the smaller Gemini margin reflecting its stronger naive baseline. Gains over the ML baselines hold for all three models. Adding indirect evidence improves mean performance in every backbone-benchmark combination, while supplying the same evidence to baselines as auxiliary multi-task targets lowers mean performance on both benchmarks. Ablations support the contributions of learned transfer scoring, informativeness, and diversity to evidence selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.