acceptodds
Under review as a conference paper at ICLR 2027

CrystalGPD: Benchmarking the Capability Boundaries of Crystal Large Language Models

Abstract

Crystal large language models (LLMs) have emerged as a flexible paradigm for crystal structure generation, yet their capabilities remain difficult to characterize due to heterogeneous tasks, interfaces, representations, datasets, and evaluation protocols. To bridge this gap, we introduce CrystalGPD (Crystal Generation, Prediction, and Design), a unified benchmark for characterizing crystal LLM capability boundaries across de novo generation, crystal structure prediction, composition- and symmetry-constrained design, and symmetry-constrained design. CrystalGPD preserves native model interfaces while standardizing task instances, generation protocols, representation-specific parsing, and downstream evaluation. Our evaluation reveals recurring limitations in jointly achieving physical feasibility and novelty, maintaining validity and structural recovery as complexity increases, satisfying exact crystallographic symmetry, and generalizing from MP20 to MP52. In particular, composition is preserved more reliably than exact symmetry, and symmetry controllability degrades for lower-frequency space groups. Motivated by these findings, we further study CrystalReAct, a benchmark-driven agentic case study that provides on-demand crystallographic knowledge and verification during generation. CrystalReAct substantially improves physical validity and crystallographic constraint satisfaction, demonstrating that external knowledge can alleviate several limitations exposed by CrystalGPD. Together, CrystalGPD provides a unified framework for evaluating crystal LLMs and, more importantly, a systematic characterization of where their current capabilities succeed and break down, providing actionable guidance for developing more reliable crystal-generation systems. Anonymous code is available at https://anonymous.4open.science/r/CrystalGPD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.