Rethinking Information Propagation in Diffusion-Based Pose-Guided Person Image Synthesis
Abstract
Pose-guided person image synthesis aims to generate realistic human images under arbitrary target poses while preserving appearance consistency. Existing diffusion-based methods have achieved remarkable progress, yet they still suffer from erroneous information propagation, where appearance features are indiscriminately propagated across semantically unrelated body regions under large pose variations. To address this issue, we propose a diffusion-based framework that explicitly rethinks information propagation during denoising. Specifically, selective self-attention suppresses redundant feature interactions through dual-sided channel gating, while the proposed query-gated cross-attention dynamically regulates appearance injection according to semantic relevance, effectively preventing erroneous feature propagation. Furthermore, a transformer refiner equipped with window-based attention and perceptual loss is introduced to further correct residual geometric inconsistencies and recover high-frequency texture details lost during latent-space compression. Extensive experiments demonstrate that our method achieves state-of-the-art generative quality, attaining competitive FID and LPIPS while requiring substantially fewer parameters and lower GPU memory.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.