ChemVerse: Benchmarking Active Learning for Virtual Screening of Million-Scale Libraries Across Docking Oracles and Biological Targets
Abstract
Virtual screening is widely used in early-stage drug discovery to prioritize promising molecules before costly experimental testing. Molecular docking enables this process at large scale, but exhaustively screening modern chemical libraries can become computationally prohibitive. Iterative, model-based selection offers an alternative by evaluating a limited subset with a docking oracle, training a surrogate model on the resulting scores, and using rapid inference to prioritize a much larger chemical space. However, it remains unclear which aspects of this setup most strongly determine the performance of adaptive virtual screening. We introduce ChemVerse, a benchmark comprising 389 million docking scores across seven protein targets and three chemical libraries ranging from molecules to molecules. Across four distinct docking oracles, seven biological targets, and three chemical libraries, complemented by publicly available DOCK3 screens, we systematically compare six molecular representations, three uncertainty-aware surrogate models, and four acquisition strategies. We find that the docking oracle strongly determines the learnability of its score landscape, producing substantially greater variation in optimization performance than molecular representation and uncertainty quantification. Three-dimensional representations perform consistently well across oracles, while uncertainty quantification has comparatively little effect. ChemVerse provides a systematic framework for understanding which components of ML-guided virtual screening generalize across oracles and which must be reconsidered as the screening environment changes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.