Guided Image Enhancement Path Diffusion Model
Abstract
Text-to-image diffusion models generate coarse structure and fine details within a single denoising process, leaving their coarse-to-fine hierarchy only implicitly modeled. We introduce GIEP (Guided Image Enhancement Path), which decomposes an image into ordered coarse-to-fine components and models their conditional dependencies. We show that the score of each partial reconstruction can be expressed as a sum of reusable per-level score contributions. This result yields a cumulative score-matching objective that supervises all levels jointly in a single shared network. Causal attention along the path axis enforces dependencies from coarse structure to fine detail, while level-specific text conditioning guides different levels of visual granularity. Experiments demonstrate leading perceptual quality and competitive compositional fidelity at a substantially smaller parameter scale; human evaluators also prefer GIEP for detail realism and saliency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.