Toward Solver-Invariant Deep Equilibrium Models
Abstract
Deep equilibrium models define their output through a fixed point of a weight-tied iteration. This formulation motivates solver invariance: predictions should be insensitive to the numerical procedure used for inference. This property need not hold at a finite computation budget. With model weights fixed, replacing the inference solver can reduce the accuracy of a trained MDEQ checkpoint to a fraction of its training-solver accuracy. We investigate the ungated identity path as a source of this sensitivity. FROST removes that bypass, scales the recurrent path with a single learned scalar, and initializes the otherwise free input scale by a self-similar rescaling rule. Without assuming a globally unique fixed point, we give sufficient conditions linking endpoint residuals, state proximity, and readout margins, and measure their premises on trained models. Under identical implicit training, FROST holds a substantially smaller accuracy spread across seven inference procedures than MDEQ. A constant input scale reproduces it, and the result transfers to a non-vision Poisson task.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.