MolTokSyn: Molecule-Text Retrieval With Text-Conditioned Tokens and Structure-Grounded Synthetic Captions
Abstract
Natural language provides a flexible way to search for molecules and describe the chemical properties that matter for a given task. However, existing molecule-text datasets rely largely on database or literature descriptions that vary widely in content and do not consistently emphasize properties that can be derived from molecular structure. We address this limitation with MolSynText, a synthetic molecule-text dataset whose descriptions are generated from structured molecular information and screened for consistency with their source. We further introduce MolTokSyn, a graph-text retrieval model that represents each molecule with multiple learned molecular tokens and uses the text query to determine how these tokens are combined for scoring, allowing different aspects of the same molecule to be emphasized for different descriptions. Across bidirectional molecule-text retrieval, MolTokSyn consistently outperforms strong existing methods on MolSynText and remains competitive on matched PubChem-derived text, while its learned molecular representations transfer strongly to downstream property-prediction tasks. Because retrieval alone does not reveal which chemical concepts are captured by the learned representations, we additionally introduce a controlled Semantic Benchmark covering structural, physicochemical, and relational properties. On this benchmark, MolSynText supervision improves MolTokSyn on most semantic categories, with particularly strong gains for descriptor-grounded properties. Together, these results show that structure-grounded synthetic descriptions can provide informative chemical supervision, especially when combined with text-conditioned molecular representations.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.