Cosine Tolerance Annealing for Neural ODE Training: A Study of Cost and Accuracy
Abstract
Neural Ordinary Differential Equations (Neural ODEs) are usually trained with adaptive solvers whose tolerances stay constant for the whole run. Loose tolerances keep the number of function evaluations (NFE) low but give less accurate solutions; tight tolerances are accurate but expensive at every iteration. We study Cosine Tolerance Annealing (CTA), which tightens the relative and absolute tolerances from loose initial values to tight final values along a cosine curve over the first 60% of training, and compare it with constant tolerances and other schedules. On a synthetic cubic system with the Dormand-Prince solver and ten seeds, CTA has on every seed a lower error, averaged over the last ten checkpoints, than each of the four constant tolerances we tested; relative to the tight final tolerance (Fixed-Final), it lowers this error by 62% and the training NFE by 21%. Across five synthetic settings, CTA needs 13-40% fewer training NFE than Fixed-Final, with an error that is lower in two settings and not significantly different in the other three; a linear schedule is not significantly different from the cosine one in four of them. On MNIST (three seeds) and Jena Climate forecasting (ten seeds), CTA reaches an accuracy similar to that of Fixed-Final with 22.8% and 17.8% fewer function evaluations per training iteration, but a constant tolerance between the initial and final values (Fixed-Geometric) reaches a similar accuracy at a lower average cost. On the two synthetic systems, however, this constant tolerance is significantly less accurate than Fixed-Final, and CTA has a lower error than it in four of the five settings, so CTA is useful when it is not known in advance whether a loose constant tolerance is accurate enough.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.