ProtGround: Functional Phrase Grounding in Local Protein 3D Structural Regions
Abstract
Protein functions often depend on local 3D environments, making it important to localize functional descriptions to their supporting structural regions. However, existing protein–text pretraining methods mainly operate at the whole-protein level, leaving these local correspondences unresolved. To this end, we introduce ProtGround, a fine-grained multimodal pretraining framework that integrates protein sequences, local protein 3D structural regions, and functional descriptions to ground functional phrases in the regions that support them. Specifically, we first construct ProtGround-Align, a large-scale dataset of fine-grained local region–function pairs. Building on this dataset, we propose ProtGround-Hier, a pretraining method that localizes the local protein regions supporting a functional phrase through joint modeling of protein–region–phrase relationships. It uses multi-positive region–phrase alignment, residue–token optimal transport, and multi-level masked modeling. Extensive experiments show that ProtGround consistently outperforms strong baselines on fine-grained grounding and protein understanding tasks. These results show that grounding functional language in local regions enables function-aware, spatially localized protein understanding and improves transfer across tasks. The data and code will be made publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.