acceptodds
Under review as a conference paper at ICLR 2027

MolLangData: A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method

Abstract

Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential for enabling large language models (LLMs) to both understand structure for chemical reasoning and generate molecules from natural-language design intent. However, the substantial cost of human annotation makes it infeasible to construct large-scale, high-quality datasets of structure-grounded descriptions. This work proposes a fully automated annotation framework for generating precise molecular descriptions at scale, such that the original molecule can be unambiguously reconstructed from the description alone. Our approach extends a rule-based chemical nomenclature parser to interpret IUPAC names and construct enriched, XML metadata that explicitly encodes molecular structure. This is then used to guide LLMs in producing accurate natural-language descriptions. Using this framework, we curate MolLangData, a dataset of approximately 163k molecule-description pairs. A rigorous validation protocol combining expert human and LLM-based reconstruction on a subset of 2,000 molecules demonstrates 98.6% description precision. Using the curated dataset, we train a 4B-parameter LLM via large-scale reinforcement learning for language-conditional molecule generation. Despite its much smaller size, the resulting model significantly outperforms most of the frontier LLMs while using far fewer tokens. The proposed framework and dataset provide a reliable foundation for molecule-language alignment, readily beneficial to broader chemical tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.