acceptodds
Under review as a conference paper at ICLR 2027

Learning a PRECISE language for small-molecule binding

Abstract

Querying a given protein target against a large set of candidate small-molecule binders is a crucial first step in the drug-discovery pipeline. With billion-compound libraries now available, accurate and fast virtual screening is critical for identifying potential drugs. Sequence-based drug-target interaction (DTI) prediction methods are fast and highly scalable, but leave important challenges in localizing binding sites, handling targets beyond single-chain proteins, and selecting candidates for expensive downstream assessment. We introduce PRECISE, which learns a structure-based protein representation and its compatibility with a quantized codebook of small-molecule space. Specifically, we build on the molecular codebook introduced by CoNCISE and formulate DTI prediction as querying quantized drug representations against a geometric neural network-based representation of the protein surface enriched with electrostatic and geometric features. PRECISE addresses key limitations of previous approaches while remaining highly accurate and scalable. It also generalizes automatically to protein complexes and metal-binding proteins without specialized training. High-throughput virtual screening often fails to align well with downstream workflows. To fix this, we introduce PRECISE-MCTS, a tree search algorithm that uses feedback from a downstream scoring oracle to refine PRECISE hits under a limited time or evaluation budget. Across Vina, Boltz, and Protenix, PRECISE-MCTS acts as a force multiplier for downstream methods. Under matched evaluation budgets, it improves top-10 mean oracle scores in 29 of 30 target–oracle comparisons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.