MatEval: A Benchmark for LLMs from Knowing to Doing in the Materials Preparation Research Loop
Abstract
Large language models (LLMs) hold promise for materials science, yet existing benchmarks fall short of covering the full research loop of real-world materials preparation. To bridge this gap, we present MatEval, a comprehensive benchmark assessing LLMs across the materials preparation research loop: from foundational competencies (Materials Knowledge QA and Instruction Understanding) through five stages (Literature Understanding, Data Retrieval, Experimental Planning, Experimental Execution, and Experimental Result Analysis). MatEval contains 24 tasks with a total of 2,135 expert-curated instances and stage-specific scoring protocols for evaluating model capabilities from knowing to doing. We evaluate 24 general-purpose and materials-specialized LLMs. Results reveal systematic performance differences across stages, with no model performing consistently across the full loop. Notably, materials-specialized models show no consistent advantage over general-purpose models. MatEval redefines evaluation from answering materials questions to operating across the materials preparation research loop, providing evidence that current LLMs remain unable to complete the full research loop.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.