acceptodds
Under review as a conference paper at ICLR 2027

GEMBench: Separating Distribution Matching from Discovery in Crystal Generation

Abstract

Generative models for crystals can serve two distinct aims: reproducing the distribution embodied in the training dataset, or discovering new stable materials beyond it. Classical generative metrics (Fr\'echet distances, kernel MMD, and related feature-space distances) evaluate only the first aim, while widely used S.U.N.-style metrics (fraction of stable, unique, and novel structures) are often reported as discovery progress without declaring the discovery aim. We make this split explicit as the *generative discovery dilemma*: the two aims share the same notion of stability, yet prescribe different optimal output distributions, so evaluation must declare the aim. We then introduce GEMBench (Generative Evaluation for Materials), which scores the two aims separately as and under explicit choices of stability probe and crystal-similarity measure. A shared multi-probe stability panel (energy above hull, phonon, and finite-temperature checks) replaces the single energy-above-hull bit used by S.U.N.-style metrics. Both scores exclude near-duplicates of training-set crystals to avoid rewarding memorization. Beyond that, rewards dataset-like distributional shapes, while rewards a large and diverse set of stable structures. On public generators trained on MP-20, , , and CHGNet S.U.N. rank the same models differently. Which generator leads depends on the declared aim and on how crystal similarity is defined. In particular, the CHGNet S.U.N. leader ranks last on , because it concentrates on a few modes of the crystal space rather than discovering many distinct crystals.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.