When Data Ordering Helps or Hurts: A Matched-Support Study of Interatomic Potentials
Abstract
Curriculum learning reorganizes training data to guide the optimization of machine-learned interatomic potentials. Its effect on joint energy and force accuracy remains unclear because reported gains often combine ordering with changes in sample selection and exposure. We isolate ordering on OC20 using matched training pools and one-pass budgets, holding initialization, within-model loss weights, and evaluation conditions fixed. Five orderings are compared across four models and four in- and out-of-distribution validation splits. These comparisons reveal that ordering changes the Energy–Force balance even under the same training objective. Difficulty-based and coverage-to-tail curricula increase mean Force MAE by 10.1–16.6% versus uniform on complete validation splits, with penalties also present under constant learning rates. Increasing the pool from 200k to 500k improves the uniform Force baseline but preserves substantial ordering costs. At 500k, coverage-first reduces PaiNN Energy MAE by 24.3% while increasing Force MAE by 0.14%. No tested curriculum jointly lowers both mean errors relative to uniform. Data ordering therefore changes the Energy–Force balance under fixed exposure and loss weights, making curriculum value depend on the prediction objective and training regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.