Visual Entity Oriented Diffusion Coding for ImaGene Compression
Abstract
Generative image editing produces collections of closely related images, each comprising a source image and its edited variants. Such collections impose growing demands on compression and storage. We call these collections **ImaGenes**, as shared content is inherited across variants while edits alter only localized regions, just as genes are inherited while mutations create variation. Within an ImaGene, edits operate on *visual entities*, identifiable units of image content. Independent encoding redundantly represents these shared entities, while joint compression exploiting their preservation and variation remains underexplored. To this end, we propose **ImaGene Compression**, a joint compression paradigm built upon visual entities as its basic units. It represents each ImaGene through shared entities and their variations. We realize this paradigm with **EnDiC** (Visual Entity Oriented Diffusion Codec), a training-free diffusion-based compression framework. EnDiC localizes entity changes, shares diffusion reconstruction trajectories for unchanged content across variants, and encodes the diffusion direction choices that reconstruct changed entities as incremental bitstreams. Furthermore, EnDiC adapts bit allocation across spatial regions and denoising steps to improve coding efficiency. We also introduce **ImaGeneSet**, a dataset for this task, covering entity addition, removal, and modification. High-fidelity generative editing with quality screening yields clearly defined entity changes while keeping unchanged regions spatially consistent with the source. Experiments on ImaGeneSet demonstrate that EnDiC achieves better performance on several full-reference perceptual metrics at low bitrates. Our code and data will be made publicly available upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.