acceptodds
Under review as a conference paper at ICLR 2027

Chemical-Support-Resolved Evaluation of Model-Selection Stability under Molecular Extrapolation

Abstract

Molecular benchmarks are commonly used not only to estimate predictive performance, but also to select which model should be deployed. This practice implicitly assumes that the model ordering observed on a benchmark remains informative when target compounds move beyond the chemical support of the training data. We examine this assumption through chemical-support-resolved evaluation, which characterizes each held-out drug by its maximum similarity to the training collection and evaluates both predictive performance and model-order stability as support decreases. Across RNA–drug, drug–protein, and drug–microbe association tasks, together with a separate ChEMBL37 direct-human drug–protein evaluation, we find that conventional cold-start partitions do not define a uniform extrapolation regime: even under Scaffold-Cold, 8.2–15.3% of test drugs retain relatively close training analogues. Predictive degradation and model-order stability are distinct. In the extreme-support regime (), the controlled model set exhibits substantial ranking transitions in Yamanishi_08, MDAD, and ChEMBL37; in Yamanishi_08, Random Pair and low-support rankings yield Spearman and 48.8% near-tie-adjusted pairwise reversals. Equal-size moderate-support controls show significantly greater ranking agreement, indicating that the transition is not explained by sparse cohort size alone. These results show that a benchmark leaderboard should be interpreted relative to the chemical-support regime in which model selection will be applied, rather than assumed to remain informative across molecular extrapolation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.