Moonwalker: Learning Accurate Inversion Steps in Diffusion Models
Abstract
Standard DDIM inversion attempts to recover the noise underlying an image, but relies on a local approximation of the reverse process whose errors accumulate across steps. As a result, the recovered latent remains entangled with information from the source image. We introduce Moonwalker, a lightweight solution that corrects this approximation by learning ”forward” inversion directly from backward generation steps. Specifically, we train a LoRA adapter to take a cleaner latent and predict the exact previous noise from the sampling trajectory, perfectly retracing the path to the correct noisier state. The adapter is activated only during inversion and deactivated during generation, leaving the original diffusion model and sampling process unchanged and incurring no inference overhead. Moonwalker reduces the correlation between recovered latents and their source images, and we evaluate its effect on downstream performance across image content, style, and music editing benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.