Ligand-conditioned Discrete Diffusion Protein Co-design Model
Abstract
Proteins perform their biological functions through three-dimensional structures encoded by amino acid sequences, and ligand-binding protein co-design requires models that generate sequence–structure compatible proteins under explicit ligand constraints. Although continuous diffusion and flow-based models have enabled ligand-aware protein design in coordinate or latent feature spaces, existing discrete diffusion protein language models mainly operate over sequence or structure tokens without direct small-molecule conditioning. We introduce ProtLiD, a Protein Ligand-conditioned Discrete Diffusion model for protein sequence–structure co-design. ProtLiD jointly generates amino-acid sequence and discrete structure tokens while incorporating ligand chemical and geometric information through geometry-aware cross-attention. Trained on over one million ligand–protein complexes, ProtLiD extends masked discrete diffusion from general sequence–structure generation to ligand-aware functional protein design. For decoding, we combine MDLM-based stochastic candidate selection with confidence-margin ranking under a geometric retention budget, committing argmax predictions at retained positions and deferring the remaining proposals. Experimentally, ProtLiD improves global fold confidence over Complexa in ligand-conditioned whole-protein design, increasing TM-score from to and pLDDT from to . In ligand-binding pocket co-design, ProtLiD reduces active-site BB-RMSD from Å for FAIR and Å for PocketGen to Å, and improves ligand-aware combined pass rates over PocketGen from to and from to under increasingly stringent docking thresholds. These results demonstrate the potential of ligand-conditioned discrete diffusion as an effective token-space framework for functional protein co-design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.