ExFiT-Bench: What Transfers When Universal Interatomic Potentials are Fine-Tuned on Experiments?
Abstract
Universal machine learning interatomic potentials (uMLIPs) are now widely adopted to enable simulations across diverse materials, which makes their evaluation increasingly important. Most benchmarks measure agreement with density functional theory (DFT), while some compare properties predicted by pretrained models with experiment. When measurements are available, however, a potential can be fine-tuned to match them before further use. We introduce ExFiT-Bench to evaluate uMLIPs in this setting. The benchmark focuses on ensemble-derived properties, for which experimental measurements provide no labels for individual configurations. Under a shared protocol, we fine-tune six uMLIPs on five tasks spanning elastic constants, a phase transition and a liquid structure factor in titanium, diamond and MgO. We evaluate the target observable and transfer to held-out temperatures, another polymorph, other observables and compositions. Fine-tuning reduces the on-target error for every model on every task, but the pretrained ranking does not consistently carry over. The correction often transfers to new conditions of the target observable, but transfer across observables and compositions is less reliable and can reduce accuracy. We release the benchmark data, code and a leaderboard for evaluating new uMLIPs and fine-tuning methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.