acceptodds
Under review as a conference paper at ICLR 2027

Beyond In-Domain Evaluation: Benchmarking Experimental Generalization in Disordered Crystal Generation

Abstract

Crystal generation models are commonly trained and evaluated on open databases that combine experimentally measured and computationally generated structures. This creates a distribution-shift blind spot: strong in-domain performance may reflect database fit rather than generalization to laboratory materials. We introduce Disorder-Bench, a multi-source benchmark for measuring experimental generalization in disordered crystal generation. It contains 91,704 structures from COD, ICSD, AMCSD, and Springer Materials, organized into positional, site, and mixed positional–site disorder. The benchmark provides a unified evaluation protocol covering structural and compositional validity, distributional fidelity, occupancy statistics, Wyckoff-site diversity, and space-group preservation. Across representative diffusion, flow-matching, equivariant, and language-model generators, models trained on COD degrade substantially when evaluated on experimentally curated sources. MatterGen reaches 69.5% overall validity, whereas CrystaLLM drops to 0.0%; mixed positional–site disorder is consistently the most challenging regime. Failure analysis identifies occupancy handling, multiplicity and space-group inconsistencies, and CIF representation errors as major sources of failure. Disorder-Bench provides a reproducible testbed for separating database fit from empirical fidelity in crystal generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.