acceptodds
Under review as a conference paper at ICLR 2027

CoAE: A Visual Tokenizer with Complementary Latent Subspaces for Image Generation

Abstract

Vision foundation models (VFMs) have emerged as effective encoders for image generation due to their rich semantic information. However, VFM features often lack low-level details, leading to inferior reconstruction quality. Existing methods attempt to address this issue by fusing semantic and detail information into a homogeneous latent space. This forces the two types of information to compete within the same representation, leading to a reconstruction–generation tradeoff. In this work, we propose Complementary AutoEncoder (CoAE), which organizes the original VFM features and detail features as distinct subspaces within a heterogeneous latent space, allowing them to preserve their own structures and serve complementary roles. We further introduce a diffusion training strategy tailored to this latent organization. Ablation studies show that homogeneous latent designs rely on alignment to improve generation quality at the cost of reconstruction, while CoAE performs best without alignment. On class-conditional ImageNet , CoAE achieves strong performance in both reconstruction and generation, with an rFID of 0.21 and a gFID of 2.16 without guidance after only 80 epochs. CoAE also transfers effectively to text-to-image generation and outperforms other latent spaces.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.