VQ-DSGAN: Mitigating Latent Representation Collapse via Dual-Space Discrimination
Abstract
Vector-quantized tokenizers are a key interface for modern visual generation, yet their effectiveness is often limited by latent representation collapse. In practice, this failure appears not only as codebook collapse, where a small subset of codewords dominates usage, but also as dimensional collapse, where quantized latents occupy a narrow low-rank subspace despite a high nominal code dimension. Existing remedies are largely quantizer-centric and do not explicitly couple balanced code usage with image-space distinguishability and latent-space recoverability. We propose VQ-DSGAN, a shared energy-based adversarial framework for standard learned-codebook VQ tokenizers that addresses these failure modes in a unified manner. VQ-DSGAN leverages a single VQ-VAE backbone in generator and discriminator modes via mode-specific feature-wise affine modulation, and induces a bidirectional dual-space discriminator from the same discriminator-mode subnetwork: an encoder-decoder energy for image-space discrimination and a decoder-encoder energy for latent-space discrimination. The image-space objective provides dense, structurally grounded supervision for perceptual fidelity, while the latent-space objective combines a code-centered prior with noise-injected latent recovery to encourage balanced code usage and locally invertible latent geometry. Experiments show that VQ-DSGAN improves reconstruction fidelity and downstream class-conditional generation while achieving near-complete codebook utilization and efficiently addressing dimensional collapse without changing the standard learned-codebook VQ formulation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.