Skipping Recomputation During Training: from Recurrent Models to Model Predictive Control and Beyond
Abstract
Sequential model inference can often be accelerated with a binary ”skip or compute” gate. At each step of the sequential computation, the gate determines if an expensive operation is performed, or if it is skipped. The setup is shared across e.g. dynamical systems, diffusion models, sequence classification, and replanning in Model Predictive Control (MPC). During training, however, current methods still perform the costly updates at every step to obtain gradient signals needed for learning the gate. We address this limitation by showing that the gate can be learned from sparse execution while still receiving dense supervision over the skipped steps, without ever having used the gate to control recomputation. The gate learning signal is derived from the computational graph of the sequence model by combining downstream loss cotangents with the input sensitivity of the expensive operation. This decouples learning of the gate network from the sequential model’s recomputation schedule, enabling efficient batch-synchronized training with sparse expensive updates. We demonstrate how this approach can be applied in three distinct scenarios: autonomous dynamics, sequence classification with input-skipping recurrent models and replanning in MPC, with substantial reductions in training time and comparable or improved accuracy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.