TrimSO(2): Fast and Memory-Efficient Training of Equivariant Interatomic Potentials
Abstract
Equivariant interatomic potentials enable accurate and efficient simulations of molecules and materials. Among them, models based on SO(2) convolutions reduce the cost of tensor products and achieve state-of-the-art accuracy, but their training remains time-consuming and memory-intensive. Existing training implementations materialize large edge tensors for representation transformations surrounding dense matrix multiplications, increasing memory usage and data movement. We observe that these transformations can be integrated into the model's existing dataflow, avoiding separate intermediate materialization. We introduce TrimSO(2), a fast and memory-efficient algorithm for SO(2)-based equivariant models. It realizes this principle through rotation on gather, recombination with activation, and inverse rotation on reduction, with support for the higher-order differentiation required for training. By avoiding intermediate materialization, TrimSO(2) delivers 2.71–16.73 operator speedups across forward, backward, and double backward and reduces model-level high-bandwidth memory traffic by 3.81–9.88. Finally, TrimSO(2) achieves 2.07–2.94 training speedups and 1.53–3.16 reductions in peak training memory over PyTorch implementations for official eSEN, UMA, and EquiformerV3 workloads on A100 and H100 GPUs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.