acceptodds
Under review as a conference paper at ICLR 2027

UniFacet: One Visual Space, Complementary Facets for Image Generation and Editing

Abstract

Unified image generation and editing require a visual representation that preserves both semantic structure and fine-grained image details. Existing unified tokenizers often fine-tune pretrained vision foundation models (VFMs) under semantic and reconstruction supervision. We introduce UniFacet, a unified visual tokenizer that reformulates this problem as completing a pretrained semantic representation with reconstruction-relevant details. UniFacet keeps the VFM frozen as a semantic anchor and learns a lightweight residual branch to recover missing image-specific information, while regularizing its magnitude to obtain the necessary reconstruction details with minimal modification to the pretrained semantic structure. A lightweight projection compresses the resulting unified representation into a compact reconstruction facet for image decoding. To provide complementary semantic information for generative modeling, we derive a semantic facet from the same unified representation and combine it with the reconstruction facet as the generative target. We instantiate UniFacet in UniFacet-1.5B, a unified model for text-to-image generation and instruction-guided image editing. The model uses the unified representation for source-image conditioning and jointly predicts the reconstruction and semantic facets for image synthesis. Experiments demonstrate high-fidelity reconstruction with largely preserved pretrained semantic capabilities, together with strong text-to-image generation and instruction-guided image editing performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.