acceptodds
Under review as a conference paper at ICLR 2027

What Makes One-Step Inference Work? A Decomposition of Gradient Origin Networks

Abstract

Gradient Origin Networks remove the encoder, instead inferring a latent code with a single gradient step through the decoder with respect to the origin. We ask what a decoder must learn for such a restricted inference rule to work. We find that autoencoder decoders often reconstruct competitively when their latent codes are optimised directly, yet perform poorly under the same one-step inference rule. The difference therefore lies in inference rather than reconstruction capacity. In the linear case this has an exact explanation: one-step inference uses the decoder transpose while optimal inference uses its pseudoinverse, with the two coinciding when the decoder Gram matrix is isotropic up to scale. An untied autoencoder need not learn this parametrisation because its encoder can absorb arbitrary invertible changes of decoder coordinates. Nonlinear GONs similarly learn flatter origin-Jacobian spectra, but imposing this geometry on an autoencoder is not sufficient. Instead, successful one-step inference depends on training reconstruction through the decoder's own inference map. A decoder can reconstruct well and still be poorly suited to gradient inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.