GlassBench: Benchmarking Large Language Models for Domain-Specific Glass Materials
Abstract
Large language models (LLMs) have shown substantial potential for scientific question answering, reasoning, and knowledge integration, yet materials science still lacks standardized benchmarks for assessing domain-specific capabilities. This limitation is especially pronounced in glass science, where terminology is specialized, composition-structure-property relations are highly nonlinear, and most available resources are structured numerical databases rather than natural-language evaluation sets. Here, we introduce GlassBench, a multitask benchmark for evaluating LLMs in glass materials science. Based on this, GMatLLM is constructed as a domain-specific LLM for the glass research. First, GlassBench was constructed by extracting question-answer pairs from authoritative textbooks and research literature, followed by multistage quality control combining LLM-based scoring, confident-learning-based screening, domain-consistency checks, and expert review. The resulting dataset contains 19,974 quality-screened samples spanning judgement, single-choice, multiple-choice, short-answer, and open-ended tasks. To test the benchmark's sensitivity to domain adaptation, we used Qwen2.5-7B-Instruct as the base model and applied supervised fine-tuning with LoRA through the LLaMA-Factory framework, yielding GMatLLM. We then evaluated general-purpose, chemistry-specific, materials-specific, and encoder-only models on GlassBench and complementary general benchmarks. GMatLLM achieved approximately 0.72 accuracy on GMat-MCQ and 0.31 on GMat-QA, while its reduced scores on several general benchmarks exposed a domain-specialization/general-capability trade-off. These results indicate that GlassBench can distinguish model families and quantify both the gains and costs associated with domain adaptation. GlassBench therefore provides a reproducible framework for evaluating and developing LLMs for glass materials research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.