Multimodal Domain-Adaptive Language Models for Materials Discovery in Metal–Organic Frameworks
Abstract
Accurate and data-efficient prediction of metal–organic framework (MOF) properties requires representations that capture both chemical and structural information. We propose a two-stage framework that continually pretrains the Granite language model on PubChem records and MOF literature, and then fuses its frozen MOFid representations with 18 interpretable CIF-derived descriptors through cross-attention. Continued pretraining alone improves QMOF band-gap and HMOF gas-adsorption prediction, and structural grounding adds further gains. Controlled ablations show that combining pretrained MOFid language representations with structural descriptors outperforms using either the pretrained MOFid representations or the structural descriptors alone, while cross-attention generally performs better than concatenation under matched conditions. Representation strengths are property-specific: structurally pretrained models perform best on QMOF and CH, whereas our multimodal framework achieves the best results on all three CO targets. Bond Connectivity descriptors are most effective for QMOF, and Pore Structure descriptors are effective to HMOF. The two modalities remain complementary in low-data settings, demonstrating that compact structural grounding can effectively enhance MOFid-based language representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.