acceptodds
Under review as a conference paper at ICLR 2027

Moonwalker: Learning Accurate Inversion Steps in Diffusion Models

Abstract

Standard DDIM inversion attempts to recover the noise underlying an image, but relies on a local approximation of the reverse process whose errors accumulate across steps. As a result, the recovered latent remains entangled with information from the source image. We introduce Moonwalker, a lightweight solution that corrects this approximation by learning ”forward” inversion directly from backward generation steps. Specifically, we train a LoRA adapter to take a cleaner latent and predict the exact previous noise from the sampling trajectory, perfectly retracing the path to the correct noisier state. The adapter is activated only during inversion and deactivated during generation, leaving the original diffusion model and sampling process unchanged and incurring no inference overhead. Moonwalker reduces the correlation between recovered latents and their source images, and we evaluate its effect on downstream performance across image content, style, and music editing benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.