Learning Discrete Representations via Mutually Exclusive Random Variables
Abstract
We propose a novel approach, named Mutually Exclusive Random Variable Autoencoder (MERVar-AE), to learn discrete representations based on Mutually Exclusive Random Variables (MERVar). Mutually Exclusive Random Variable (MERVar) arises from the observation that deterministic mappings (e.g., neural decoders) involving random variables (noise injection) tend to reduce overlap coefficients among involved random variables, thereby inducing exclusivity. This exclusivity can be utilized to learn binary latents. Specifically, by injecting symmetric noise into the output of a bounded activation function (e.g., ), the output values are pushed toward the boundaries ( or for ), effectively producing binary latent representations. Compared to previous approaches based on Straight-Through Estimator (STE), such as VQ-VAE, the proposed approach is simpler, with only an activation and noise injection operation. The forward and backward computations are aligned, as gradients are propagated through the same continuous latent representation. Comprehensive experiments on nine standard datasets show that the proposed approach achieves competitive results compared to state-of-the-art methods in discrete representation learning. Controlled experiments further demonstrate the benefit of avoiding surrogate-gradient mismatch during encoder optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.