XRDBench: Benchmarking AI Agents on Experimental X-ray Diffraction Analysis
Abstract
Frontier multimodal large language models (MLLMs) increasingly demonstrate strong scientific-reasoning and coding capabilities, but whether these capabilities translate into reliable execution of domain-specific experimental workflows remains unclear. We investigate MLLM agents for X-ray diffraction (XRD) analysis, a widely used materials-characterization technique whose interpretation remains labor-intensive and expertise-dependent. We introduce XRDBench, a benchmark of 14 tasks organized into four tiers spanning peak-level perception, scientific unit analyses, multi-step characterization, and complex workflows such as phase identification, quantitative phase analysis, and Rietveld refinement. The tasks use predominantly expert-labeled experimental thin-film and powder data and are executed through a sandboxed agentic harness that provides multimodal inputs and access to scientific and crystallographic tools. We evaluate five proprietary frontier models and one open-weight model using task-specific outcome metrics and seven-axis process rubrics. We find that frontier models demonstrate strong XRD knowledge, although no single frontier model consistently outperforms the others, while the evaluated open-weight model remains behind. Removing visual access generally degrades performance, emphasizing the multimodal nature of XRD analysis. Agentic analysis may be particularly valuable for high-throughput and autonomous laboratories, where jointly analyzing related patterns across a compositional gradient reduces token use per pattern without accuracy loss. These findings highlight both the current limitations and the potential of agentic materials characterization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.