Learning the Graph Language of Molecules: Grammar Tokenization for Chemical LLMs
Abstract
Large language models (LLMs) have been explored for integrating molecular graphs, but joint autoregressive modeling of text and molecules remains challenging. A key barrier is the lack of principled molecular graph tokenizers (MGTs): string representations like SMILES rely on brittle syntax and canonicalization, limiting robustness, flexibility, and interpretability. We introduce AlphaGrammar, a neural-symbolic MGT that rewrites molecular graphs as linear derivations over a formal graph grammar with theoretical guarantees. Training optimizes the parameters of a potential function for grammar construction, and inference amortizes the cost of parsing during tokenization. We then introduce Gramole, a multimodal LLM built on AlphaGrammar, which interleaves text and molecules as rule sequences without the need for modality-specific adapters. Gramole offers % validity via type-constrained decoding, expressivity by sampling high-diversity molecular sets, transferability by improving FDA drugs' predicted properties on all six endpoints under an independent property predictor (vs. four for the strongest baseline), with edits that stay closer to the parent drug, controllability using natural language interfaces, and synthesizability by achieving drug and material retrosynthesis success rate of the best multimodal LLM baseline at matched Qwen3-8B backbone. AlphaGrammar autonomously acquires rules that recapitulate known functional groups, offering a layer of interpretability absent in prior MGTs. We benchmark Gramole across diverse tasks in the design-make-test-analysis (DMTA) cycle and probe AlphaGrammar's mechanism in quantitative and qualitative case studies.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.