Risk-Controlled DFT Selection for Machine-Learned Interatomic Potentials
Abstract
Density functional theory (DFT) gives accurate energies and forces but is expensive. Pretrained machine-learned interatomic potentials (MLIPs) are widely used for their cheap costs, but their errors vary widely across molecules. So how do we choose which configurations to trust with MLIP and which to send to DFT? We cast this as risk-controlled selection and solve it with a conformal risk control (CRC) framework. A finite CRC rule turns any error score fixed before calibration into a routing decision, whose expected capped per-molecule error left to the MLIP is at most a user-chosen fraction of its fitting-set mean. Because every score receives the same guarantee, scores can be compared by the DFT cost or call. With this comparison across eight dataset–MLIP pairs, we find that committee disagreement often overlooks errors that the models share. We also show that an analytical-chemistry-informed lightweight error predictor ranks errors more accurately than existing error-score baselines. At it sends 7% and 14% fewer than two committee-based baselines while trailing by 3.4 percentage points from oracle. The advantage is emphasized when the user's DFT reference differs from the MLIPs' training reference and is less apparent when the references match. We also show that replacing only some MLIP energies with DFT can break error cancellation in energy differences, so guarantees on relative energies must be calibrated directly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.