The Materials Scientist Agent: Reinforcement Learning for Crystal Discovery
Abstract
Discovering new inorganic crystals can be seen as a search problem, where the evaluation reveals if a candidate is stable, but not how to improve it. For this purpose, we develop an LLM-based agent and train a 4B model to do targeted materials search. We give the model computational tools, have it propose candidates in a Wyckoff representation, check their novelty, and evaluate their stability and properties over multiple turns. To stop the agent from resubmitting one good material, we score a group of attempts together and split credit between repeated structures. We evaluate on in-domain compositions from MP-20, out-of-domain from WBM, and band-gap targets, judging every structure with MLIPs not used in training. Reinforcement learning nearly doubles the rate of stable, unique, and novel (S.U.N.) structures over supervised fine-tuning, on known and unknown compositions, and the resulting 4B agent outperforms GLM-5.3 Flash and nearly matches Gemini 3.1 Pro, the teacher model. For band-gap targets, it finds up to more S.U.N. wide-gap materials than a state-of-the-art language model for crystal generation. We also show how the agent can be used to search for lead-free perovskites for solar cells.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.