GCMatch: Protein Pocket Matching with Geometric–Chemical Spatial Semantic Modeling and Fine-Grained Representations
Abstract
Protein pocket similarity matching supports drug repurposing and pocket function annotation. This task requires similarity in both cavity geometry and chemical environment, i.e., dual similarity. Existing learning-based methods typically encode geometric and chemical information jointly within the same local unit. They produce single-vector pocket representations for matching, enabling effective assessment of overall similarity. However, they do not explicitly model spatial semantic relations between cavity geometry and chemical environment, and the coarse-grained single-vector representation obscures local geometric or chemical differences. Moreover, the datasets used by these methods define similarity labels according to whether pockets bind similar ligands, but ligand similarity is not equivalent to dual similarity. Together, these issues limit the accuracy of dual-similarity pocket matching. We propose GCMatch, a dual-similarity pocket matching method based on spatial semantic modeling and fine-grained representations. First, GCMatch extracts cavity geometric points and chemical functional points from each pocket. It then connects both point types in a single neighborhood relation graph. A graph encoder learns two sets of fine-grained node representations, one for geometry and one for chemistry. Multi-vector matching is applied separately to each set, and product fusion produces the final matching score. GCMatch achieves a test AUC of 0.920, exceeding the best baseline by 5.1 percentage points. We also construct a four-quadrant dataset with separate geometric and chemical similarity labels for each pocket pair. It serves as a unified evaluation benchmark for the field.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.