Efficient and Faithful High-Resolution Image Editing with Sparse Cross-Resolution Denoising
Abstract
Instruction-based image editors remain predominantly 1K systems, yet practical editing increasingly demands faithful 2K/4K outputs. Direct high-resolution adaptation preserves a unified editing trajectory but incurs dense target–source attention over rapidly growing token sequences; low-resolution editing followed by external refinement is cheaper but can separate detail restoration from the editing trajectory. We introduce sparse cross-resolution denoising, a native high-resolution editing paradigm that retains a frozen 1K stream as a semantic and structural anchor and trains a target-resolution stream to reconstruct local detail within the same Transformer. An asymmetric sparse topology preserves full 1K attention, sends spatially aligned noisy-target guidance from the low-resolution trajectory to high-resolution queries, and restricts high-resolution target–reference interaction to local Halo neighborhoods. Resolution-specific diffusion-state mapping and spatial YaRN adapt the editor to larger grids. A training-free progressive path further replaces an initial portion of target-resolution denoising with a 1K prefix and completes refinement in the same model. On VINS-4KEval, we achieve the highest reported metrics at both 2K and 4K; in a matched 4K forward benchmark, the sparse topology is \(4.2\times\) faster overall and \(25.3\times\) faster in attention than dense attention. Progressive inference further reduces end-to-end latency by \(30.5%\) at 4K while retaining competitive quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.