TopU-LBVS: A Realistic Multi Target Benchmark for Ligand Based Virtual Screening
Abstract
Ligand-based virtual screening (LBVS) is a practical first-pass tool in early-stage drug discovery, but current benchmarks often overestimate progress through random negatives, easy decoys, limited target coverage, and non-standardized evaluation protocols. We introduce TopU-LBVS, a multi-target benchmark for LBVS under hard-negative screening conditions. Starting from curated ChEMBL 35 bioactivity data, TopU-LBVS covers 93 protein targets across 7 protein classes and constructs target-specific screening libraries with property-matched, structurally similar decoys at a fixed 1:40 active-to-decoy ratio. Libraries contain roughly 400 to 10,000 compounds and are designed to reduce simple physicochemical and nearest-neighbor fingerprint shortcuts. TopU-LBVS provides three fixed protocols. TopU-LBVS-full evaluates ChEMBL* -> TopU generalization across all 93 targets. TopU-LBVS-few evaluates few-shot TopU -> TopU learning, where both training and test compounds come from hard-negative libraries. TopU-LBVS-mini gives a compact seven-target protocol with a paired random-decoy control, enabling low-cost development and direct measurement of the gap between random ChEMBL* decoys and TopU hard decoys. Across ten reference baselines, including classical fingerprint methods, molecular GNNs, fingerprint hybrids, and modern molecular models, TopU-LBVS shows that performance on standard random-decoy evaluations can degrade sharply under hard-negative screening. We release the data, fixed splits, evaluation code, and baseline implementations to support reproducible comparison of future LBVS methods and molecular foundation models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.