Replay the Seeds, Store the Remainder: Compressing Persistent Diffusion Inversion State
Abstract
Inversion-based diffusion editors retain dense stochastic controls to support repeated text-guided edits from a single inversion, imposing severe storage and bandwidth bottlenecks. We observe that a substantial component of each materialized control is deterministically governed by the inversion solver and its generating random draws. Leveraging this insight, we introduce *replay factorization*, a denoiser-free representation that reconstructs an analytical predictor directly from recorded provenance, projects the remainder onto solver-prescribed stochastic carrier directions, and compactly codes the residual defect alongside endpoint and metadata. We formally establish the *stencil-consistency condition* under which explicit source-latent terms vanish, derive exact replay coordinates for first- and second-order stochastic solvers, and generalize the construction to shifted paths and optimized noise. Evaluated on all 700 PIE-Bench tasks across four U-Net configurations (Edit-Friendly and LEDITS++), replay factorization achieves 35.77–47.04% median complete-file savings over optimized generic codecs while preserving edit fidelity (median LPIPS 0.00023–0.0023 to raw-state edits). Without per-source fitting, parameter-free analytical predictors retain the majority of these gains. On the PixArt- diffusion Transformer, our lossy format saves 45.25% storage, while an exact-state variant achieves bitwise input restoration with 14.21% median savings over native-lossless baselines. The exact framework similarly compresses four-step ReNoise and TurboEdit states without loss. By transforming generating provenance into reproducible structure at the decoder, replay factorization substantially reduces persistent editing state without evaluating neural denoisers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.