QuantaMind: A Large Chemistry Language Model for Structured Molecular Screening and Discovery
Abstract
Large language models offer a flexible interface for scientific reasoning, yet general-purpose models remain poorly calibrated for quantitative molecular prediction, especially when targets depend on excited-state electronic structure. We introduce QuantaMind, a 32-billion-parameter chemistry language model that maps a SMILES string and a chemistry instruction to a structured, multi-property molecular profile for screening and discovery. QuantaMind is obtained by parameter-efficient instruction tuning of Qwen-2.5-32B on QuantumChem-200K, a supervision corpus of more than 214,000 organic molecules annotated with photophysical, quantum-chemical, safety, accessibility, and physicochemical properties. The model is evaluated on 3,000 unseen molecules from a hold-out testset using both property-prediction and decision-oriented screening metrics. QuantaMind achieves an overall weighted mean absolute error (wMAE) of 0.1980, compared with 3.3040 for its untuned Qwen backbone and 0.5297 for a Gemma-3-27B model fine-tuned under the same protocol. In multi-objective screening, it recovers 35 of the ground-truth top 100 candidates, reaches 10.50-fold enrichment, and obtains 3.26% normalized hypervolume regret. Data-scaling, training-duration, backbone-transfer, and failure-mode analyses further characterize where domain adaptation helps and where SMILES-only prediction remains limited. These results position QuantaMind as a chemistry language model for quantitatively grounded molecular screening, with photoinitiator discovery as a demanding case study.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.