acceptodds
Under review as a conference paper at ICLR 2027

PLI-CLIP: Learning Transferable Protein–Ligand Interaction Representations from Natural Language Supervision

Abstract

Predicting the 3D structure of protein–ligand complexes is central to structure-based drug discovery, and recent approaches employ protein language models (PLMs) to address this challenge. While PLMs provide powerful representations of proteins, they are inherently ligand-agnostic and do not capture explicit protein–ligand interaction semantics. To address this shortcoming, we exploit rich, text-like biochemical annotations produced by tools such as PLIP to learn ligand-aware protein representations. Specifically, we construct the Protein-Ligand Interaction Text Dataset (PLI-TD) by translating structured PLIP profiles into grounded, natural-language interaction descriptions. We then introduce PLI-CLIP, a CLIP-style contrastive framework that aligns protein features with PLI-TD to inject explicit interaction semantics. While these textual descriptions specify which interactions (e.g., hydrogen bonds, hydrophobic contacts) occur, they do not localize them. Therefore, PLI-CLIP is further trained to predict per-residue binding geometry, transforming geometry-blind representations into geometry-aware ones. Finally, these representations serve as a transferable module that integrates into existing structure prediction models without modifying their original architectures, eliminating the need for text annotations or interaction profiles at inference time. PLI-CLIP consistently improves docking accuracy across various recent downstream docking models, raising DiffDock's apo top-1 rate (24.4 → 30.3), NeuralPLexer's <2 Å fraction (36.6 → 40.2), and FlowDock's docking success rate (50.4 → 54.1).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.