Pixel Art Is a Palette, Not a Downscale: Rebuilding Targets and Decoding for Sprite Generation
Abstract
Game sprites of 12–24 px are among the smallest images people draw on purpose, and their medium is defined by a handful of flat colours and hard alpha edges. Despite rapid progress in text-to-image generation, existing systems render such sprites as shrunken illustrations with soft shading and around a hundred colours. We identify two root causes: the training targets of sprite corpora are dominated by downscaled artwork rather than native pixel art, and continuous diffusion has no mechanism to express the few-colour structure of the medium. To address these challenges, we propose PaletteDiff, a framework for native-resolution sprite generation. First, we introduce Native-Resolution Target Rebuilding, which removes duplicate leakage, forbids aggressive downscaling and adds 15k permissively licensed sprites, together with a native evaluation protocol that scores against real low-resolution pixel art only. Second, we propose Palette-Projected Sampling, a training-free decoding rule that projects the predicted clean image onto a per-sprite k-colour palette during the late denoising steps, so that the remaining steps repair the quantisation seams. Third, combined with SDEdit-style re-noising, the same rule yields a Palette Refinement module for sprites from generators we did not train. PaletteDiff lowers FD-DINOv2 at 16 px from 167.2 to 31.8 ± 0.5; a per-sprite palette is the decisive ingredient, and imposing it inside the sampler improves on the same quantisation applied afterwards by a further 18%. Under the native protocol, PaletteDiff has the lowest FD among gpt-image-2, FLUX.2 and SDXL pipelines given the same palette post-process; Palette Refinement improves all three external systems by 36–41% while preserving 80% of their silhouettes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.