UniFacet: One Visual Space, Complementary Facets for Image Generation and Editing
Abstract
Unified image generation and editing require a visual representation that preserves both semantic structure and fine-grained image details. Existing unified tokenizers often fine-tune pretrained vision foundation models (VFMs) under semantic and reconstruction supervision. We introduce UniFacet, a unified visual tokenizer that reformulates this problem as completing a pretrained semantic representation with reconstruction-relevant details. UniFacet keeps the VFM frozen as a semantic anchor and learns a lightweight residual branch to recover missing image-specific information, while regularizing its magnitude to obtain the necessary reconstruction details with minimal modification to the pretrained semantic structure. A lightweight projection compresses the resulting unified representation into a compact reconstruction facet for image decoding. To provide complementary semantic information for generative modeling, we derive a semantic facet from the same unified representation and combine it with the reconstruction facet as the generative target. We instantiate UniFacet in UniFacet-1.5B, a unified model for text-to-image generation and instruction-guided image editing. The model uses the unified representation for source-image conditioning and jointly predicts the reconstruction and semantic facets for image synthesis. Experiments demonstrate high-fidelity reconstruction with largely preserved pretrained semantic capabilities, together with strong text-to-image generation and instruction-guided image editing performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.